fix(report): stop "not distinguishable" reading as a pass - #13
Merged
Merged
Conversation
The inconclusive verdict said "not distinguishable from a perfectly calibrated model at this sample size". That is exactly what the arithmetic establishes and close to the opposite of what it means. A reader skimming a report saw their model compared against a perfect one, no difference found, and took it as a clean result. What it actually means is that the dataset is too small to resolve the question either way: a well calibrated model and a badly calibrated one both land inside the band on too few rows, and the figure does not say which you have. The tool's whole argument is that an unreadable number should not look like a finding, and its own headline verdict was doing precisely that. Both call sites now lead with INCONCLUSIVE and say plainly that nothing was established in either direction. The report's "How to read this" preamble defines the term rather than assuming it, because the report is what users read, not the README. Two tests now pin the valence rather than the wording: one asserts the statement contains "not a result in either direction" and contains no pass-like phrasing, and one asserts the preamble carries the definition. The committed example report is regenerated, its header rewritten to match, and the README, METHODOLOGY and PLAN no longer use the old phrasing. The README's quotations were verbatim and are verbatim again. Also corrects "every figure is read against its own null" to "every inferential figure": cost and latency are measurements rather than inferences and correctly have no null. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The inconclusive verdict read as a pass.
That is exactly what the arithmetic establishes, and close to the opposite of
what it means. A reader skimming sees their model compared against a perfect
one, no difference found, and takes it as a clean result. It actually means the
dataset is too small to resolve the question either way: a well calibrated
model and a badly calibrated one both land inside the band on too few rows, and
the figure does not say which you have.
This is the tool's own headline verdict doing the exact thing the tool exists
to prevent, which is why it was worth its own pull request.
After
Fixed in the report output first, not just the README, because the report
is what users read. Both call sites changed: the calibration verdict and the
baseline figure statement. The report's "How to read this" preamble now defines
the term instead of assuming it.
Tests pin the valence, not the wording
test_an_inconclusive_statement_does_not_read_as_a_passasserts thestatement says "not a result in either direction" and contains none of
passes,acceptable,looks fine,no issue,is calibrated.test_the_preamble_says_inconclusive_is_not_a_passasserts the reportcarries the definition.
Four existing tests asserted the old string and were updated to assert the new
meaning rather than the new string.
Knock-on
docs/example-report.mdregenerated so the committed page shows the newwording, with its hand-written header rewritten to match. The README's four
block quotes were verbatim and are verbatim again, verified
whitespace-normalised against the regenerated report. METHODOLOGY and PLAN no
longer use the old phrasing anywhere.
Also corrects "every figure is read against its own null" to "every
inferential figure" — cost and latency are measurements, not inferences,
and correctly have no null. That was a literal overstatement.