-
Notifications
You must be signed in to change notification settings - Fork 34
Add optional dataset_description and column_descriptions to Task #191
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -340,6 +340,11 @@ class Task: | |
| name of 2 parent directories for local or S3-based datasets. | ||
|
|
||
| This field is only here for convenience and is not used for any validation when computing the results. | ||
| dataset_description : str | None, default None | ||
| Text description of the dataset. | ||
| column_descriptions : dict[str, str] | None, default None | ||
| Text description of each column used by the task. If provided, the keys must exactly match the target, | ||
| dynamic and static columns of the task. | ||
|
|
||
| Examples | ||
| -------- | ||
|
|
@@ -379,6 +384,8 @@ class Task: | |
| past_dynamic_columns: list[str] = dataclasses.field(default_factory=list) | ||
| static_columns: list[str] = dataclasses.field(default_factory=list) | ||
| task_name: str | None = None | ||
| dataset_description: str | None = None | ||
| column_descriptions: dict[str, str] | None = None | ||
|
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. This will also be part of evaluation summaries. Are we okay with that?
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. The summaries can end up being really large. What do you think about only storing the hash/fingerprint of |
||
|
|
||
| def __post_init__(self): | ||
| if self.task_name is None: | ||
|
|
@@ -458,6 +465,16 @@ def __post_init__(self): | |
| "`generate_univariate_targets_from` cannot be used for multivariate tasks (when `target` is a list)" | ||
| ) | ||
|
|
||
| if self.column_descriptions is not None: | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Let's also raise if |
||
| expected = set(self.target_columns + self.dynamic_columns + self.static_columns) | ||
| missing = sorted(expected - set(self.column_descriptions)) | ||
| unexpected = sorted(set(self.column_descriptions) - expected) | ||
| if missing or unexpected: | ||
| raise ValueError( | ||
| "`column_descriptions` must have exactly one entry per target, dynamic and static column of the " | ||
| f"task. Missing: {missing}. Unexpected: {unexpected}." | ||
| ) | ||
|
|
||
| # Attributes computed after the dataset is loaded | ||
| self._full_dataset: datasets.Dataset | None = None | ||
| self._freq: str | None = None | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Let's maybe call this
task_description? I think that fits a bit better since it's a property of the task and multiple descriptions can be given to the same dataset