Skip to content

fix(loaders): refuse what cannot be scored honestly, in both loaders and the pricing file - #86

Merged
TMHSDigital merged 1 commit into
mainfrom
fix/loader-validation
Sep 23, 2026
Merged

TMHSDigital merged 1 commit into
mainfrom
fix/loader-validation

Conversation

@TMHSDigital

Copy link
Copy Markdown
Owner

Fixes #44. Fixes #45.

#44: load_jevbench skipped the duplicate-id check that load_jsonl makes, so ["dup", "dup"] loaded with no refusal. Both loaders now share one check and one message.

#45, dataset loader:

  • Byte order mark: a UTF-8 BOM (which Windows editors and spreadsheet exports write) made line 1 unparseable. The first line is now read as utf-8-sig.
  • Not UTF-8: a non-UTF-8 file failed with a bare codec error. Lines are now decoded one by one, and the error names the file, the line and the byte.
  • Labels: str() turned a label of null, true or 1 into "None", "True" or "1" and loaded an option nobody wrote, and empty or blank labels were accepted. Every label must now be a non-empty string. The public fixture's labels are all strings, and it still loads 111 of 111.
  • Descriptions: label_descriptions could describe options that don't exist, and null became "None". Keys must now be options, and values non-empty strings.

#45, pricing loader:

  • NaN, infinite and negative prices were accepted. They are now refused.
  • as_of accepted 20260901 and 2026-W36-1 (fromisoformat takes both) and dates in the future. A future date would keep the report from ever calling a price stale. All three are now refused.
  • The pricing file is read as utf-8-sig too.

The "Your own data" page lists the new refusals.

Checked

  • New tests failed on main and pass now:
    • a duplicate JevBench id is refused;
    • a BOM file loads;
    • a latin-1 line is named ("latin1.jsonl ... line 2");
    • five kinds of bad label are each refused;
    • three kinds of bad description are each refused;
    • six bad pricing values are each refused;
    • a BOM pricing file loads.
  • ruff, mypy --strict, the full pytest suite, and the site build all pass. The example report is unchanged.

🤖 Generated with Claude Code

…and the pricing file

#44: load_jevbench skipped the duplicate-id check load_jsonl makes, so
["dup", "dup"] loaded with no refusal and later joins by id could mix
rows. Both loaders now share one check and one message.

#45, datasets:
- A UTF-8 byte order mark, which Windows editors and spreadsheet exports
  write, made line 1 unparseable; the first line is read as utf-8-sig.
- A file that is not UTF-8 failed with a bare codec error. Lines are
  decoded one by one, and the error names the file, the line, and the
  byte.
- str() turned a label of null, true or 1 into "None", "True" or "1" and
  loaded an option nobody wrote; empty and blank labels were accepted.
  Every label must be a non-empty string.
- label_descriptions could describe options that do not exist, and null
  became "None". Keys must be options and values non-empty strings.

#45, pricing: NaN, infinite and negative prices were accepted; as_of took
20260901 and 2026-W36-1 (fromisoformat accepts both) and future dates,
which would keep the report from ever calling a price stale. Each is
refused, and the pricing file is read as utf-8-sig too.

The "Your own data" page lists the new refusals.

Fixes #44. Fixes #45.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@TMHSDigital
TMHSDigital merged commit cedd1cb into main Sep 23, 2026
17 checks passed
@TMHSDigital
TMHSDigital deleted the fix/loader-validation branch September 23, 2026 20:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Input validation gaps in the dataset and pricing loaders load_jevbench skips the duplicate-id check

1 participant