Add models.yaml as the single source of truth for leaderboard metadata - #190
Merged
Merged
Conversation
Submitters now declare how their model is presented on the fev-bench leaderboard (organization, link, model type, zero-shot, commercial use) in benchmarks/fev_bench/models.yaml, next to their results CSV, instead of a maintainer filling in a dict in the leaderboard app afterwards. - One model per results CSV, so a models.yaml key (the file name) identifies exactly one leaderboard row. Splits toto-2_0.csv into one file per checkpoint size. - scripts/fev_bench_metadata.py validates the schema and both directions of the file <-> entry mapping; the submission tests and the leaderboard app both read metadata through it. - model_type adds `system` and `closed-api` values for submissions whose results cannot be reproduced independently or compared to single models. - `hidden: true` replaces the hardcoded Toto exclusion list in the figure script (and the leaderboard's own copy of it).
|
Check out this pull request on See visual diffs & provide feedback on Jupyter Notebooks. Powered by ReviewNB |
Most pairwise skill scores between the top models are within ±10%, so with a ±30% range half of the heatmap cells were barely tinted.
- ModelMetadata (scripts/fev_bench_metadata.py) replaces the hand-written schema checks and is the single reference for the models.yaml format. models.yaml, models/README.md, the tutorial and evaluate.py now point to it instead of repeating the field list. - Drop the tests that only exercised the schema checks; the submission test still loads all metadata and checks it against the results files. - Fix the pretrained / task-specific type descriptions.
shchur
marked this pull request as ready for review
October 2, 2026 07:53
shchur
commented
Oct 2, 2026
Hidden models are now looked up in models.yaml, which describes exactly the files in benchmarks/fev_bench/results, so a custom results directory failed the one-to-one metadata check. To include an unsubmitted model, add its CSV and models.yaml entry to the results directory.
abdulfatir
approved these changes
Oct 2, 2026
abdulfatir
left a comment
Collaborator
There was a problem hiding this comment.
Thanks Oleks! Only minor questions/comments.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Submitters now declare how their model appears on the fev-bench leaderboard in
benchmarks/fev_bench/models.yaml, in the same PR as their results CSV. Previously a maintainer filled inMODEL_CONFIGin the leaderboard Space by hand. After this, updating the leaderboard is a single command in the Space:uv run python save_tables.py [commit].Changes
One model per results CSV. The
models.yamlkey is the results CSV file name, so each key points to exactly one leaderboard row.toto-2_0.csvis split into one file per checkpoint size; the rows are unchanged (checked withassert_frame_equal).scripts/fev_bench_metadata.pydefines the format as a pydantic model,ModelMetadata, and checks that entries and results files match one-to-one. It is the only place the format is documented;models.yaml,models/README.md, the tutorial andevaluate.pypoint to it. The submission tests and the leaderboard Space both load metadata through this module.model_typegainssystemandclosed-api. The leaderboard hides models of these types until the reader opts in.closed-apitakes precedence oversystem.hidden: truereplaces the hardcoded Toto exclusion list ingenerate_fev_bench_figures.py.Tests (
test_fev_bench_submissions.py): one new test loads all metadata, which validates every entry and the one-to-one match with the results files.evaluate.pypoints to the submission instructions after a run.Docs: submission steps updated in
models/README.mdand the add-your-model tutorial.Pairwise skill score heatmap color range narrowed from ±30% to ±15% (separate commit), since most pairwise skill scores are within ±10%. This changes the paper figures.
Review notes
commercial_use: falsefor TiRex, TimesFM-3, TabPFN-TS-3, CITRAS-FM, and Moirai-2.0 (its HF card says CC-BY-NC-4.0). All others aretrue. TS-ICL and TiRex-2 use custom licenses I haven't read; worth confirming.lingjiang2_apihasurl: TODO. Needs a real link before merge.models.yamlthroughfev_bench_metadata.py, so it depends on this PR.Testing
uv run pytest test/test_fev_bench_submissions.py: 36 passed.test_fev_bench_submissions.py,test_analysis.py,test_model.py,test_task.py,test_utils.py(168 passed). Not run locally:test_metrics.py(needsautogluon) andtest_notebooks.py(needsnbformat), because those packages aren't installed here.generate_fev_bench_figures.pyandrender_fev_bench_leaderboard.pyrun cleanly; the figure script excludes the 3 hidden Toto sizes.save_tables.py --fev-repo ..succeeds, and the app has no errors with toggles on or off.