Skip to content

Add models.yaml as the single source of truth for leaderboard metadata - #190

Merged
shchur merged 5 commits into
mainfrom
model-metadata
Oct 2, 2026
Merged

shchur merged 5 commits into
mainfrom
model-metadata

Conversation

@shchur

@shchur shchur commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Submitters now declare how their model appears on the fev-bench leaderboard in benchmarks/fev_bench/models.yaml, in the same PR as their results CSV. Previously a maintainer filled in MODEL_CONFIG in the leaderboard Space by hand. After this, updating the leaderboard is a single command in the Space: uv run python save_tables.py [commit].

my-model:                    # = results CSV file name
  organization: My Org
  url: https://huggingface.co/my-org/my-model
  model_type: pretrained     # pretrained | task-specific | statistical | system | closed-api
  zero_shot: true
  commercial_use: true       # required, no default
  # display_name, hidden: optional

Changes

  • One model per results CSV. The models.yaml key is the results CSV file name, so each key points to exactly one leaderboard row. toto-2_0.csv is split into one file per checkpoint size; the rows are unchanged (checked with assert_frame_equal).

  • scripts/fev_bench_metadata.py defines the format as a pydantic model, ModelMetadata, and checks that entries and results files match one-to-one. It is the only place the format is documented; models.yaml, models/README.md, the tutorial and evaluate.py point to it. The submission tests and the leaderboard Space both load metadata through this module.

  • model_type gains system and closed-api. The leaderboard hides models of these types until the reader opts in. closed-api takes precedence over system.

  • hidden: true replaces the hardcoded Toto exclusion list in generate_fev_bench_figures.py.

  • Tests (test_fev_bench_submissions.py): one new test loads all metadata, which validates every entry and the one-to-one match with the results files.

  • evaluate.py points to the submission instructions after a run.

  • Docs: submission steps updated in models/README.md and the add-your-model tutorial.

  • Pairwise skill score heatmap color range narrowed from ±30% to ±15% (separate commit), since most pairwise skill scores are within ±10%. This changes the paper figures.

Review notes

  • Licenses: commercial_use: false for TiRex, TimesFM-3, TabPFN-TS-3, CITRAS-FM, and Moirai-2.0 (its HF card says CC-BY-NC-4.0). All others are true. TS-ICL and TiRex-2 use custom licenses I haven't read; worth confirming.
  • lingjiang2_api has url: TODO. Needs a real link before merge.
  • The matching leaderboard Space change is committed locally and not yet pushed. It reads models.yaml through fev_bench_metadata.py, so it depends on this PR.

Testing

  • uv run pytest test/test_fev_bench_submissions.py: 36 passed.
  • Ran these suites locally: test_fev_bench_submissions.py, test_analysis.py, test_model.py, test_task.py, test_utils.py (168 passed). Not run locally: test_metrics.py (needs autogluon) and test_notebooks.py (needs nbformat), because those packages aren't installed here.
  • generate_fev_bench_figures.py and render_fev_bench_leaderboard.py run cleanly; the figure script excludes the 3 hidden Toto sizes.
  • Leaderboard Space run against this branch: save_tables.py --fev-repo .. succeeds, and the app has no errors with toggles on or off.

Submitters now declare how their model is presented on the fev-bench
leaderboard (organization, link, model type, zero-shot, commercial use)
in benchmarks/fev_bench/models.yaml, next to their results CSV, instead
of a maintainer filling in a dict in the leaderboard app afterwards.

- One model per results CSV, so a models.yaml key (the file name)
  identifies exactly one leaderboard row. Splits toto-2_0.csv into one
  file per checkpoint size.
- scripts/fev_bench_metadata.py validates the schema and both directions
  of the file <-> entry mapping; the submission tests and the leaderboard
  app both read metadata through it.
- model_type adds `system` and `closed-api` values for submissions whose
  results cannot be reproduced independently or compared to single models.
- `hidden: true` replaces the hardcoded Toto exclusion list in the figure
  script (and the leaderboard's own copy of it).
@review-notebook-app

Copy link
Copy Markdown

Check out this pull request on  ReviewNB

See visual diffs & provide feedback on Jupyter Notebooks.


Powered by ReviewNB

shchur added 2 commits October 2, 2026 07:47
Most pairwise skill scores between the top models are within ±10%, so with
a ±30% range half of the heatmap cells were barely tinted.
- ModelMetadata (scripts/fev_bench_metadata.py) replaces the hand-written
  schema checks and is the single reference for the models.yaml format.
  models.yaml, models/README.md, the tutorial and evaluate.py now point to
  it instead of repeating the field list.
- Drop the tests that only exercised the schema checks; the submission
  test still loads all metadata and checks it against the results files.
- Fix the pretrained / task-specific type descriptions.
@shchur
shchur marked this pull request as ready for review October 2, 2026 07:53
@shchur
shchur requested a review from abdulfatir October 2, 2026 07:53
Comment thread scripts/fev_bench_metadata.py
Hidden models are now looked up in models.yaml, which describes exactly the
files in benchmarks/fev_bench/results, so a custom results directory failed
the one-to-one metadata check. To include an unsubmitted model, add its CSV
and models.yaml entry to the results directory.

@abdulfatir abdulfatir left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks Oleks! Only minor questions/comments.

Comment thread benchmarks/fev_bench/models.yaml Outdated
Comment thread benchmarks/fev_bench/models.yaml
Comment thread models/README.md
Comment thread scripts/fev_bench_metadata.py
@shchur
shchur merged commit 510551b into main Oct 2, 2026
6 checks passed
@shchur
shchur deleted the model-metadata branch October 2, 2026 10:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants