Add optional dataset_description and column_descriptions to Task - #191
abdulfatir wants to merge 3 commits into
Conversation
Both default to None. When column_descriptions is provided, its keys must exactly match the task's target, dynamic and static columns.
| static_columns: list[str] = dataclasses.field(default_factory=list) | ||
| task_name: str | None = None | ||
| dataset_description: str | None = None | ||
| column_descriptions: dict[str, str] | None = None |
There was a problem hiding this comment.
This will also be part of evaluation summaries. Are we okay with that?
There was a problem hiding this comment.
The summaries can end up being really large. What do you think about only storing the hash/fingerprint of task_description and column_descriptions in the evaluations summary? This will ensure reproducibility (hash doesn't match -> tasks scores are not comparable) without blowing up the size of the CSV? Alternatively we can just switch to storing/loading summaries as parquet (which doesn't work nicely with git though)
| name of 2 parent directories for local or S3-based datasets. | ||
|
|
||
| This field is only here for convenience and is not used for any validation when computing the results. | ||
| dataset_description : str | None, default None |
There was a problem hiding this comment.
Let's maybe call this task_description? I think that fits a bit better since it's a property of the task and multiple descriptions can be given to the same dataset
| static_columns: list[str] = dataclasses.field(default_factory=list) | ||
| task_name: str | None = None | ||
| dataset_description: str | None = None | ||
| column_descriptions: dict[str, str] | None = None |
There was a problem hiding this comment.
The summaries can end up being really large. What do you think about only storing the hash/fingerprint of task_description and column_descriptions in the evaluations summary? This will ensure reproducibility (hash doesn't match -> tasks scores are not comparable) without blowing up the size of the CSV? Alternatively we can just switch to storing/loading summaries as parquet (which doesn't work nicely with git though)
| "`generate_univariate_targets_from` cannot be used for multivariate tasks (when `target` is a list)" | ||
| ) | ||
|
|
||
| if self.column_descriptions is not None: |
There was a problem hiding this comment.
Let's also raise if self.column_descriptions is not None and generate_univariate_targets_from is not None? Otherwise the descriptions become ambiguous
Adds two optional fields to
Task:dataset_description: str | None = None: text description of the dataset.column_descriptions: dict[str, str] | None = None: text description of each column used by the task. If provided, its keys must exactly match the task's target, dynamic and static columns; otherwise a validation error lists the missing / unexpected columns.Both default to
None, so existing tasks and benchmark YAMLs are unaffected.