Skip to content

fix(baseline): read accuracy and AUROC nulls in both directions - #79

Merged
TMHSDigital merged 1 commit into
mainfrom
fix/nulls-test-both-tails
Sep 23, 2026
Merged

TMHSDigital merged 1 commit into
mainfrom
fix/nulls-test-both-tails

Conversation

@TMHSDigital

Copy link
Copy Markdown
Owner

Fixes #33.

The problem. Accuracy and AUROC were compared only with the upper tail of their null. A model reliably worse than the null fell far below it and was reported as INCONCLUSIVE, with the advice to collect more rows. An inverted score over 300 rows printed AUROC 0.0000 ... INCONCLUSIVE ... Collect more rows: the wrong conclusion, and the wrong advice.

What changed

  • NullBand now carries a 5th percentile (p05).
  • A value below it gets its own verdict, with the likely reason and the percentile it fell below:
    • accuracy: "worse than chance", because the answers are systematically wrong, usually from misaligned labels and options;
    • AUROC: "ranks incorrect above correct", because the score is inverted.
  • Above the 95th percentile, or between the two, a value reads exactly as before. No existing report line changes, and the site build confirms the example report still matches its command line for line.
  • METHODOLOGY now says both of these nulls are read in both directions, and why the calibration floors stay one-sided.

Checked

  • New tests failed on main and pass now:
    • an inverted score (AUROC 0.0 over 300 rows) is below the null, and its statement names the 5th percentile and says "inverted", not INCONCLUSIVE;
    • all-wrong accuracy on four-way rows reads "worse than chance";
    • 25% accuracy on four-way rows is still INCONCLUSIVE.
  • ruff, mypy --strict, the full pytest suite, and the site build all pass.

🤖 Generated with Claude Code

Accuracy and AUROC were compared only with the upper tail of their null.
A model reliably worse than the null landed far below it and was reported
as INCONCLUSIVE with the advice to collect more rows; an inverted score
over 300 rows printed "AUROC 0.0000 ... INCONCLUSIVE ... Collect more
rows", which is the wrong conclusion and the wrong advice.

NullBand now carries a 5th percentile. A value below it gets its own
verdict: accuracy "worse than chance" (the answers are systematically
wrong, usually misaligned labels and options), AUROC "ranks incorrect
above correct" (the score is inverted). Values above the 95th percentile
or between the two read exactly as before, so no existing report line
changes; the example report still matches its command line for line.
METHODOLOGY says which tails are read, and why the calibration floors
stay one-sided.

Fixes #33.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@TMHSDigital
TMHSDigital merged commit 2e17c3e into main Sep 23, 2026
17 checks passed
@TMHSDigital
TMHSDigital deleted the fix/nulls-test-both-tails branch September 23, 2026 18:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Accuracy and AUROC nulls are one-sided, so an inverted model reads as "collect more rows"

1 participant