diff --git a/METHODOLOGY.md b/METHODOLOGY.md index cb3dc67..30af29f 100644 --- a/METHODOLOGY.md +++ b/METHODOLOGY.md @@ -447,9 +447,9 @@ bootstrap draws: Read the middle row. A model whose reported probabilities have been sharpened by a temperature of 0.5 is overconfident in a way any user would notice, ECE catches it at more than five times its floor, and MCE lands at 0.98 times its -own floor: inside the band, not distinguishable from a perfectly calibrated -model. Even at T = 0.35, where ECE is nine times its floor, MCE clears its floor -by seven percent. +own floor: inside the band, so MCE reports the question as unanswerable on this +many rows while ECE answers it plainly. Even at T = 0.35, where ECE is nine +times its floor, MCE clears its floor by seven percent. That is why MCE is demoted to a diagnostics block and never printed beside ECE. It is not wrong, and it is not useless on larger samples, but at the few hundred diff --git a/README.md b/README.md index 6437af3..c96ab69 100644 --- a/README.md +++ b/README.md @@ -40,15 +40,23 @@ noise and sample size. If you do not know that number, you cannot read your own. This is not a rounding concern. On a few hundred rows a calibration claim is frequently not measurable at all. -So plumbline reports every figure against its own null, and says plainly when a -value is indistinguishable from a calibrated model. From the example report, 105 -rows: +So plumbline reports every inferential figure against its own null, and says +plainly when a value sits inside what that null already produces. + +**When it says a figure is inconclusive, that is not a pass.** It means the +dataset cannot tell your model apart from the null, so nothing was established +in either direction. A model that is genuinely well calibrated and one that is +badly calibrated can both land there on too few rows, and the figure does not +say which you have. Reading it as a clean bill of health is the single easiest +mistake to make with this tool, and it inverts the conclusion. + +From the example report, 105 rows: - ECE 0.0740, against a calibrated-model floor of 0.0707 and a 95th percentile of - 0.1109. Not distinguishable. -- Brier 0.1711, against a floor of 0.1489. Not distinguishable. + 0.1109. Inconclusive. +- Brier 0.1711, against a floor of 0.1489. Inconclusive. - Confidence AUROC 0.6034, against a permutation null of 0.4997 and a 95th - percentile of 0.6116. Not distinguishable. + percentile of 0.6116. Inconclusive. - Accuracy 0.7714, against a chance null of 0.3416. Better than chance. Four figures, one of which supports a conclusion. A tool that printed the first @@ -109,9 +117,11 @@ uv run plumbline version lines from it: > ECE 0.0740 over 105 rows (10 equal width bins), against a calibrated-model -> floor of 0.0707 (95th percentile 0.1109): not distinguishable from a perfectly -> calibrated model at this sample size. Collect more rows before reading anything -> into it. +> floor of 0.0707 (95th percentile 0.1109): INCONCLUSIVE at this sample size. A +> perfectly calibrated model would often score this badly on this many rows, so +> this dataset cannot tell the two apart. This is not a clean bill of health: +> nothing was established either way. Collect more rows to make the question +> answerable. > No cost available. None of the 105 cases could be priced, so cost is not > reported rather than being shown as zero. 105 rows: tokens were reported, but diff --git a/docs/PLAN.md b/docs/PLAN.md index 718f965..417ee77 100644 --- a/docs/PLAN.md +++ b/docs/PLAN.md @@ -310,11 +310,12 @@ rewrite. ### What the run says about the tool Accuracy 0.6500 over 40 rows against a chance null of 0.2571, better than -chance. ECE 0.1070 against a calibrated-model floor of 0.1400, **not -distinguishable from a perfectly calibrated model at this sample size**. Brier -0.1737 against a floor of 0.1512, also not distinguishable. Confidence AUROC -0.7981 against a permutation null of 0.5013, separates correct from incorrect. -Recalibration refused: it needs 200 held-out rows and this split has 20. +chance. ECE 0.1070 against a calibrated-model floor of 0.1400, **inconclusive: +40 rows cannot tell this apart from a perfectly calibrated model, which +establishes nothing in either direction**. Brier 0.1737 against a floor of +0.1512, also inconclusive. Confidence AUROC 0.7981 against a permutation null +of 0.5013, separates correct from incorrect. Recalibration refused: it needs +200 held-out rows and this split has 20. That is the intended behaviour at n = 40 and it is the argument the README makes. diff --git a/docs/example-report.md b/docs/example-report.md index 5c205e6..7572da9 100644 --- a/docs/example-report.md +++ b/docs/example-report.md @@ -21,10 +21,12 @@ which does not depend on who answered. What to look at, in order: - **ECE is 0.0740 and the floor is 0.0707.** The tool reports the figure and then - says it is not distinguishable from a perfectly calibrated model at 105 rows. - A tool that printed 0.0740 on its own would be handing you a number that looks - like a finding and is not one. Every figure on the page is read against its own - null this way. + calls it inconclusive: a perfectly calibrated model would often score this + badly on 105 rows, so these rows cannot tell the two apart. Read that as the + absence of a result, not as a pass. A tool that printed 0.0740 on its own + would be handing you a number that looks like a finding and is not one. Every + inferential figure on the page is read against its own null this way; cost and + latency are measurements rather than inferences and have no null. - **Cost says why it is blank rather than showing a zero.** The basis is "the model that answered is not priced", which is a different fact from "this adapter cannot report tokens". A zero in that column would read as free. @@ -44,12 +46,12 @@ Running the same command against a real vendor produces the same shape. See # plumbline report -Generated 2026-09-21 against dataset `c18e9496`, 105 rows. 1 arm(s). +Generated 2026-09-22 against dataset `c18e9496`, 105 rows. 1 arm(s). ## How to read this - Every figure states the rows it was computed on and the null it is read against. A number on its own is not a finding. -- "Not distinguishable" means the value sits inside what the null produces at this sample size. It does not mean the systems are the same; it means this dataset cannot tell them apart yet. +- **"INCONCLUSIVE" is not a pass.** It means the value sits inside what the null already produces at this sample size, so this dataset cannot tell the two apart. Nothing was established in either direction. A model that is genuinely well calibrated and one that is badly calibrated can both land here on too few rows, and the figure does not say which you have. - Arms are grouped by what kind of number they report. **Figures in different groups are not comparable** and are never placed side by side. - "Not reported" is a result with a reason attached, not a missing cell. @@ -71,9 +73,9 @@ The vendor asserts these probabilities are calibrated. Whether that survives con - **Asked** — 67 choice asked as choice, 38 noul asked as choice. - 38 noul rows were asked as choice questions, which is a different question from the one the dataset states. Not comparable with an arm that asked them as noul. - Accuracy 0.7714 over 105 rows, against a chance null of 0.3416 (95th percentile 0.4190): better than chance at this sample size. -- ECE 0.0740 over 105 rows (10 equal width bins), against a calibrated-model floor of 0.0707 (95th percentile 0.1109): not distinguishable from a perfectly calibrated model at this sample size. Collect more rows before reading anything into it. -- Brier 0.1711 over 105 rows, against a calibrated-model floor of 0.1489 (95th percentile 0.1839): not distinguishable from a perfectly calibrated model at this sample size. Collect more rows before reading anything into it. -- **Confidence** — AUROC 0.6034 over 105 rows, against a permutation null of 0.4997 (95th percentile 0.6116): not distinguishable from permutation at this sample size. Collect more rows before reading anything into it. +- ECE 0.0740 over 105 rows (10 equal width bins), against a calibrated-model floor of 0.0707 (95th percentile 0.1109): INCONCLUSIVE at this sample size. A perfectly calibrated model would often score this badly on this many rows, so this dataset cannot tell the two apart. This is not a clean bill of health: nothing was established either way. Collect more rows to make the question answerable. +- Brier 0.1711 over 105 rows, against a calibrated-model floor of 0.1489 (95th percentile 0.1839): INCONCLUSIVE at this sample size. A perfectly calibrated model would often score this badly on this many rows, so this dataset cannot tell the two apart. This is not a clean bill of health: nothing was established either way. Collect more rows to make the question answerable. +- **Confidence** — AUROC 0.6034 over 105 rows, against a permutation null of 0.4997 (95th percentile 0.6116): INCONCLUSIVE at this sample size. Permutation would often score this well on this many rows, so this dataset cannot tell the two apart. This is not a result in either direction. Collect more rows to make the question answerable. - **Cost** — not reported. No cost available. None of the 105 cases could be priced, so cost is not reported rather than being shown as zero. 105 rows: tokens were reported, but the model that answered is not priced. - **Latency** — p50 37.7ms, p95 85.0ms, p99 104.7ms over 105 live calls @@ -89,5 +91,5 @@ The vendor asserts these probabilities are calibrated. Whether that survives con Read these only after the figures above. MCE is a maximum over bins, decided by one bin, and at a few hundred rows it cannot detect overconfidence spread evenly across the range. -- MCE 0.1588 over 105 rows (10 equal width bins), against a calibrated-model floor of 0.1251 (95th percentile 0.2198): not distinguishable from a perfectly calibrated model at this sample size. Collect more rows before reading anything into it. -- Multiclass Brier 0.3526 over 105 rows, against a calibrated-model floor of 0.3168 (95th percentile 0.3975): not distinguishable from a perfectly calibrated model at this sample size. Collect more rows before reading anything into it. +- MCE 0.1588 over 105 rows (10 equal width bins), against a calibrated-model floor of 0.1251 (95th percentile 0.2198): INCONCLUSIVE at this sample size. A perfectly calibrated model would often score this badly on this many rows, so this dataset cannot tell the two apart. This is not a clean bill of health: nothing was established either way. Collect more rows to make the question answerable. +- Multiclass Brier 0.3526 over 105 rows, against a calibrated-model floor of 0.3168 (95th percentile 0.3975): INCONCLUSIVE at this sample size. A perfectly calibrated model would often score this badly on this many rows, so this dataset cannot tell the two apart. This is not a clean bill of health: nothing was established either way. Collect more rows to make the question answerable. diff --git a/src/plumbline/metrics/baseline.py b/src/plumbline/metrics/baseline.py index fb6a1e7..1ecf482 100644 --- a/src/plumbline/metrics/baseline.py +++ b/src/plumbline/metrics/baseline.py @@ -64,12 +64,18 @@ def is_distinguishable(self) -> bool: return self.value > self.band.p95 def statement(self) -> str: + # The inconclusive wording names the measurement as what failed, not + # the model as what passed. See _judgment in metrics/calibration.py for + # why: "not distinguishable from " is read as a pass, and it + # means nothing was established either way. judgment = ( f"{self.beats} at this sample size." if self.is_distinguishable else ( - f"not distinguishable from {self.band.null} at this sample size. " - "Collect more rows before reading anything into it." + f"INCONCLUSIVE at this sample size. {self.band.null.capitalize()} " + f"would often score this well on this many rows, so this dataset " + "cannot tell the two apart. This is not a result in either " + "direction. Collect more rows to make the question answerable." ) ) name = METRIC_NAMES.get(self.metric, self.metric.upper()) diff --git a/src/plumbline/metrics/calibration.py b/src/plumbline/metrics/calibration.py index cd94a3d..254cd8b 100644 --- a/src/plumbline/metrics/calibration.py +++ b/src/plumbline/metrics/calibration.py @@ -597,10 +597,22 @@ def verdict(measured: float, floor: FloorBand) -> str: def _judgment(measured: float, floor: FloorBand) -> str: - """The trailing clause of a verdict, so one wording serves every caller.""" + """The trailing clause of a verdict, so one wording serves every caller. + + The inconclusive wording puts the failure on the measurement, not on the + model. "Not distinguishable from a perfectly calibrated model" is literally + what the arithmetic says, and it reads as a pass: a reader skimming a + report sees their model compared to a perfect one and no difference found. + What it actually means is that this dataset is too small to resolve the + question either way, which is the absence of a result rather than a good + one. The sentence has to say so, because the sentence is what gets read. + """ if is_distinguishable(measured, floor): return "miscalibration is distinguishable from sampling noise." return ( - "not distinguishable from a perfectly calibrated model at this sample size. " - "Collect more rows before reading anything into it." + "INCONCLUSIVE at this sample size. A perfectly calibrated model would " + "often score this badly on this many rows, so this dataset cannot tell " + "the two apart. This is not a clean bill of health: nothing was " + "established either way. Collect more rows to make the question " + "answerable." ) diff --git a/src/plumbline/report/markdown.py b/src/plumbline/report/markdown.py index c6be107..bb44f38 100644 --- a/src/plumbline/report/markdown.py +++ b/src/plumbline/report/markdown.py @@ -148,9 +148,11 @@ def _how_to_read() -> list[str]: "", "- Every figure states the rows it was computed on and the null it is read " "against. A number on its own is not a finding.", - '- "Not distinguishable" means the value sits inside what the null produces at ' - "this sample size. It does not mean the systems are the same; it means this " - "dataset cannot tell them apart yet.", + '- **"INCONCLUSIVE" is not a pass.** It means the value sits inside what the ' + "null already produces at this sample size, so this dataset cannot tell the two " + "apart. Nothing was established in either direction. A model that is genuinely " + "well calibrated and one that is badly calibrated can both land here on too few " + "rows, and the figure does not say which you have.", "- Arms are grouped by what kind of number they report. **Figures in different " "groups are not comparable** and are never placed side by side.", '- "Not reported" is a result with a reason attached, not a missing cell.', diff --git a/tests/test_baseline.py b/tests/test_baseline.py index 5b536ce..06f5e7b 100644 --- a/tests/test_baseline.py +++ b/tests/test_baseline.py @@ -61,7 +61,25 @@ def test_an_accuracy_at_chance_is_not_dressed_up_as_a_result() -> None: figure = baseline.accuracy_figure([True] * 25 + [False] * 75, [4] * 100, n_boot=N_BOOT) assert not figure.is_distinguishable - assert "not distinguishable" in figure.statement() + statement = figure.statement() + assert "INCONCLUSIVE" in statement + assert "cannot tell the two apart" in statement + + +def test_an_inconclusive_statement_does_not_read_as_a_pass() -> None: + """The wording must not let a reader take "no difference found" as good news. + + The old wording was "not distinguishable from at this sample size", + which is what the arithmetic says and the opposite of what it means. A + reader skimming a report saw their model compared against a null and no + difference found, and read it as a clean result. + """ + figure = baseline.accuracy_figure([True] * 25 + [False] * 75, [4] * 100, n_boot=N_BOOT) + statement = figure.statement().lower() + + assert "not a result in either direction" in statement + for pass_like in ("passes", "acceptable", "looks fine", "no issue", "is calibrated"): + assert pass_like not in statement def test_auroc_of_an_uninformative_score_is_not_distinguishable_from_the_null() -> None: diff --git a/tests/test_calibration.py b/tests/test_calibration.py index 91849cc..66bc2df 100644 --- a/tests/test_calibration.py +++ b/tests/test_calibration.py @@ -98,7 +98,11 @@ def test_a_calibrated_mock_lands_inside_its_own_noise_floor(n_cases: int) -> Non measured_ece = calibration.ece(run.probabilities, run.outcomes) assert not calibration.is_distinguishable(measured_ece, floors["ece"]) - assert "not distinguishable" in calibration.verdict(measured_ece, floors["ece"]) + verdict = calibration.verdict(measured_ece, floors["ece"]) + assert "INCONCLUSIVE" in verdict + # The valence has to be explicit. Landing on the floor is the absence of a + # result, not a passing grade, and the sentence is what gets read. + assert "not a clean bill of health" in verdict @pytest.mark.parametrize("n_cases", [500, 4000]) diff --git a/tests/test_report.py b/tests/test_report.py index ba61e4e..1f77cb9 100644 --- a/tests/test_report.py +++ b/tests/test_report.py @@ -145,7 +145,15 @@ def test_calibration_is_read_against_its_floor() -> None: line = next(line for line in text.splitlines() if "ECE" in line) assert "floor" in line - assert "distinguishable" in line + assert "INCONCLUSIVE" in line or "distinguishable from sampling noise" in line + + +def test_the_preamble_says_inconclusive_is_not_a_pass() -> None: + """The report defines its own terms, because the report is what gets read.""" + text = render(a_run()) + + assert '**"INCONCLUSIVE" is not a pass.**' in text + assert "Nothing was established in either direction." in text def test_the_maximum_error_is_a_diagnostic_and_not_a_headline() -> None: