Skip to content

Add optional dataset_description and column_descriptions to Task - #191

Open
abdulfatir wants to merge 3 commits into
mainfrom
task-descriptions
Open

abdulfatir wants to merge 3 commits into
mainfrom
task-descriptions

Conversation

@abdulfatir

@abdulfatir abdulfatir commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Adds two optional fields to Task:

  • dataset_description: str | None = None: text description of the dataset.
  • column_descriptions: dict[str, str] | None = None: text description of each column used by the task. If provided, its keys must exactly match the task's target, dynamic and static columns; otherwise a validation error lists the missing / unexpected columns.

Both default to None, so existing tasks and benchmark YAMLs are unaffected.

Abdul Fatir Ansari added 3 commits October 2, 2026 15:08
Both default to None. When column_descriptions is provided, its keys must exactly
match the task's target, dynamic and static columns.
@abdulfatir
abdulfatir requested a review from shchur October 2, 2026 15:18
Comment thread src/fev/task.py
static_columns: list[str] = dataclasses.field(default_factory=list)
task_name: str | None = None
dataset_description: str | None = None
column_descriptions: dict[str, str] | None = None

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will also be part of evaluation summaries. Are we okay with that?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The summaries can end up being really large. What do you think about only storing the hash/fingerprint of task_description and column_descriptions in the evaluations summary? This will ensure reproducibility (hash doesn't match -> tasks scores are not comparable) without blowing up the size of the CSV? Alternatively we can just switch to storing/loading summaries as parquet (which doesn't work nicely with git though)

Comment thread src/fev/task.py
name of 2 parent directories for local or S3-based datasets.

This field is only here for convenience and is not used for any validation when computing the results.
dataset_description : str | None, default None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's maybe call this task_description? I think that fits a bit better since it's a property of the task and multiple descriptions can be given to the same dataset

Comment thread src/fev/task.py
static_columns: list[str] = dataclasses.field(default_factory=list)
task_name: str | None = None
dataset_description: str | None = None
column_descriptions: dict[str, str] | None = None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The summaries can end up being really large. What do you think about only storing the hash/fingerprint of task_description and column_descriptions in the evaluations summary? This will ensure reproducibility (hash doesn't match -> tasks scores are not comparable) without blowing up the size of the CSV? Alternatively we can just switch to storing/loading summaries as parquet (which doesn't work nicely with git though)

Comment thread src/fev/task.py
"`generate_univariate_targets_from` cannot be used for multivariate tasks (when `target` is a list)"
)

if self.column_descriptions is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's also raise if self.column_descriptions is not None and generate_univariate_targets_from is not None? Otherwise the descriptions become ambiguous

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants