fix(baseline): read accuracy and AUROC nulls in both directions - #79
Merged
Merged
Conversation
Accuracy and AUROC were compared only with the upper tail of their null. A model reliably worse than the null landed far below it and was reported as INCONCLUSIVE with the advice to collect more rows; an inverted score over 300 rows printed "AUROC 0.0000 ... INCONCLUSIVE ... Collect more rows", which is the wrong conclusion and the wrong advice. NullBand now carries a 5th percentile. A value below it gets its own verdict: accuracy "worse than chance" (the answers are systematically wrong, usually misaligned labels and options), AUROC "ranks incorrect above correct" (the score is inverted). Values above the 95th percentile or between the two read exactly as before, so no existing report line changes; the example report still matches its command line for line. METHODOLOGY says which tails are read, and why the calibration floors stay one-sided. Fixes #33. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #33.
The problem. Accuracy and AUROC were compared only with the upper tail of their null. A model reliably worse than the null fell far below it and was reported as INCONCLUSIVE, with the advice to collect more rows. An inverted score over 300 rows printed
AUROC 0.0000 ... INCONCLUSIVE ... Collect more rows: the wrong conclusion, and the wrong advice.What changed
NullBandnow carries a 5th percentile (p05).Checked
mainand pass now:ruff,mypy --strict, the fullpytestsuite, and the site build all pass.🤖 Generated with Claude Code