Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions METHODOLOGY.md
Original file line number Diff line number Diff line change
Expand Up @@ -447,9 +447,9 @@ bootstrap draws:
Read the middle row. A model whose reported probabilities have been sharpened by
a temperature of 0.5 is overconfident in a way any user would notice, ECE
catches it at more than five times its floor, and MCE lands at 0.98 times its
own floor: inside the band, not distinguishable from a perfectly calibrated
model. Even at T = 0.35, where ECE is nine times its floor, MCE clears its floor
by seven percent.
own floor: inside the band, so MCE reports the question as unanswerable on this
many rows while ECE answers it plainly. Even at T = 0.35, where ECE is nine
times its floor, MCE clears its floor by seven percent.

That is why MCE is demoted to a diagnostics block and never printed beside ECE.
It is not wrong, and it is not useless on larger samples, but at the few hundred
Expand Down
28 changes: 19 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,15 +40,23 @@ noise and sample size. If you do not know that number, you cannot read your own.
This is not a rounding concern. On a few hundred rows a calibration claim is
frequently not measurable at all.

So plumbline reports every figure against its own null, and says plainly when a
value is indistinguishable from a calibrated model. From the example report, 105
rows:
So plumbline reports every inferential figure against its own null, and says
plainly when a value sits inside what that null already produces.

**When it says a figure is inconclusive, that is not a pass.** It means the
dataset cannot tell your model apart from the null, so nothing was established
in either direction. A model that is genuinely well calibrated and one that is
badly calibrated can both land there on too few rows, and the figure does not
say which you have. Reading it as a clean bill of health is the single easiest
mistake to make with this tool, and it inverts the conclusion.

From the example report, 105 rows:

- ECE 0.0740, against a calibrated-model floor of 0.0707 and a 95th percentile of
0.1109. Not distinguishable.
- Brier 0.1711, against a floor of 0.1489. Not distinguishable.
0.1109. Inconclusive.
- Brier 0.1711, against a floor of 0.1489. Inconclusive.
- Confidence AUROC 0.6034, against a permutation null of 0.4997 and a 95th
percentile of 0.6116. Not distinguishable.
percentile of 0.6116. Inconclusive.
- Accuracy 0.7714, against a chance null of 0.3416. Better than chance.

Four figures, one of which supports a conclusion. A tool that printed the first
Expand Down Expand Up @@ -109,9 +117,11 @@ uv run plumbline version
lines from it:

> ECE 0.0740 over 105 rows (10 equal width bins), against a calibrated-model
> floor of 0.0707 (95th percentile 0.1109): not distinguishable from a perfectly
> calibrated model at this sample size. Collect more rows before reading anything
> into it.
> floor of 0.0707 (95th percentile 0.1109): INCONCLUSIVE at this sample size. A
> perfectly calibrated model would often score this badly on this many rows, so
> this dataset cannot tell the two apart. This is not a clean bill of health:
> nothing was established either way. Collect more rows to make the question
> answerable.

> No cost available. None of the 105 cases could be priced, so cost is not
> reported rather than being shown as zero. 105 rows: tokens were reported, but
Expand Down
11 changes: 6 additions & 5 deletions docs/PLAN.md
Original file line number Diff line number Diff line change
Expand Up @@ -310,11 +310,12 @@ rewrite.
### What the run says about the tool

Accuracy 0.6500 over 40 rows against a chance null of 0.2571, better than
chance. ECE 0.1070 against a calibrated-model floor of 0.1400, **not
distinguishable from a perfectly calibrated model at this sample size**. Brier
0.1737 against a floor of 0.1512, also not distinguishable. Confidence AUROC
0.7981 against a permutation null of 0.5013, separates correct from incorrect.
Recalibration refused: it needs 200 held-out rows and this split has 20.
chance. ECE 0.1070 against a calibrated-model floor of 0.1400, **inconclusive:
40 rows cannot tell this apart from a perfectly calibrated model, which
establishes nothing in either direction**. Brier 0.1737 against a floor of
0.1512, also inconclusive. Confidence AUROC 0.7981 against a permutation null
of 0.5013, separates correct from incorrect. Recalibration refused: it needs
200 held-out rows and this split has 20.

That is the intended behaviour at n = 40 and it is the argument the README
makes.
Expand Down
24 changes: 13 additions & 11 deletions docs/example-report.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,10 +21,12 @@ which does not depend on who answered.
What to look at, in order:

- **ECE is 0.0740 and the floor is 0.0707.** The tool reports the figure and then
says it is not distinguishable from a perfectly calibrated model at 105 rows.
A tool that printed 0.0740 on its own would be handing you a number that looks
like a finding and is not one. Every figure on the page is read against its own
null this way.
calls it inconclusive: a perfectly calibrated model would often score this
badly on 105 rows, so these rows cannot tell the two apart. Read that as the
absence of a result, not as a pass. A tool that printed 0.0740 on its own
would be handing you a number that looks like a finding and is not one. Every
inferential figure on the page is read against its own null this way; cost and
latency are measurements rather than inferences and have no null.
- **Cost says why it is blank rather than showing a zero.** The basis is "the
model that answered is not priced", which is a different fact from "this
adapter cannot report tokens". A zero in that column would read as free.
Expand All @@ -44,12 +46,12 @@ Running the same command against a real vendor produces the same shape. See

# plumbline report

Generated 2026-09-21 against dataset `c18e9496`, 105 rows. 1 arm(s).
Generated 2026-09-22 against dataset `c18e9496`, 105 rows. 1 arm(s).

## How to read this

- Every figure states the rows it was computed on and the null it is read against. A number on its own is not a finding.
- "Not distinguishable" means the value sits inside what the null produces at this sample size. It does not mean the systems are the same; it means this dataset cannot tell them apart yet.
- **"INCONCLUSIVE" is not a pass.** It means the value sits inside what the null already produces at this sample size, so this dataset cannot tell the two apart. Nothing was established in either direction. A model that is genuinely well calibrated and one that is badly calibrated can both land here on too few rows, and the figure does not say which you have.
- Arms are grouped by what kind of number they report. **Figures in different groups are not comparable** and are never placed side by side.
- "Not reported" is a result with a reason attached, not a missing cell.

Expand All @@ -71,9 +73,9 @@ The vendor asserts these probabilities are calibrated. Whether that survives con
- **Asked** — 67 choice asked as choice, 38 noul asked as choice.
- 38 noul rows were asked as choice questions, which is a different question from the one the dataset states. Not comparable with an arm that asked them as noul.
- Accuracy 0.7714 over 105 rows, against a chance null of 0.3416 (95th percentile 0.4190): better than chance at this sample size.
- ECE 0.0740 over 105 rows (10 equal width bins), against a calibrated-model floor of 0.0707 (95th percentile 0.1109): not distinguishable from a perfectly calibrated model at this sample size. Collect more rows before reading anything into it.
- Brier 0.1711 over 105 rows, against a calibrated-model floor of 0.1489 (95th percentile 0.1839): not distinguishable from a perfectly calibrated model at this sample size. Collect more rows before reading anything into it.
- **Confidence** — AUROC 0.6034 over 105 rows, against a permutation null of 0.4997 (95th percentile 0.6116): not distinguishable from permutation at this sample size. Collect more rows before reading anything into it.
- ECE 0.0740 over 105 rows (10 equal width bins), against a calibrated-model floor of 0.0707 (95th percentile 0.1109): INCONCLUSIVE at this sample size. A perfectly calibrated model would often score this badly on this many rows, so this dataset cannot tell the two apart. This is not a clean bill of health: nothing was established either way. Collect more rows to make the question answerable.
- Brier 0.1711 over 105 rows, against a calibrated-model floor of 0.1489 (95th percentile 0.1839): INCONCLUSIVE at this sample size. A perfectly calibrated model would often score this badly on this many rows, so this dataset cannot tell the two apart. This is not a clean bill of health: nothing was established either way. Collect more rows to make the question answerable.
- **Confidence** — AUROC 0.6034 over 105 rows, against a permutation null of 0.4997 (95th percentile 0.6116): INCONCLUSIVE at this sample size. Permutation would often score this well on this many rows, so this dataset cannot tell the two apart. This is not a result in either direction. Collect more rows to make the question answerable.
- **Cost** — not reported. No cost available. None of the 105 cases could be priced, so cost is not reported rather than being shown as zero. 105 rows: tokens were reported, but the model that answered is not priced.
- **Latency** — p50 37.7ms, p95 85.0ms, p99 104.7ms over 105 live calls

Expand All @@ -89,5 +91,5 @@ The vendor asserts these probabilities are calibrated. Whether that survives con

Read these only after the figures above. MCE is a maximum over bins, decided by one bin, and at a few hundred rows it cannot detect overconfidence spread evenly across the range.

- MCE 0.1588 over 105 rows (10 equal width bins), against a calibrated-model floor of 0.1251 (95th percentile 0.2198): not distinguishable from a perfectly calibrated model at this sample size. Collect more rows before reading anything into it.
- Multiclass Brier 0.3526 over 105 rows, against a calibrated-model floor of 0.3168 (95th percentile 0.3975): not distinguishable from a perfectly calibrated model at this sample size. Collect more rows before reading anything into it.
- MCE 0.1588 over 105 rows (10 equal width bins), against a calibrated-model floor of 0.1251 (95th percentile 0.2198): INCONCLUSIVE at this sample size. A perfectly calibrated model would often score this badly on this many rows, so this dataset cannot tell the two apart. This is not a clean bill of health: nothing was established either way. Collect more rows to make the question answerable.
- Multiclass Brier 0.3526 over 105 rows, against a calibrated-model floor of 0.3168 (95th percentile 0.3975): INCONCLUSIVE at this sample size. A perfectly calibrated model would often score this badly on this many rows, so this dataset cannot tell the two apart. This is not a clean bill of health: nothing was established either way. Collect more rows to make the question answerable.
10 changes: 8 additions & 2 deletions src/plumbline/metrics/baseline.py
Original file line number Diff line number Diff line change
Expand Up @@ -64,12 +64,18 @@ def is_distinguishable(self) -> bool:
return self.value > self.band.p95

def statement(self) -> str:
# The inconclusive wording names the measurement as what failed, not
# the model as what passed. See _judgment in metrics/calibration.py for
# why: "not distinguishable from <null>" is read as a pass, and it
# means nothing was established either way.
judgment = (
f"{self.beats} at this sample size."
if self.is_distinguishable
else (
f"not distinguishable from {self.band.null} at this sample size. "
"Collect more rows before reading anything into it."
f"INCONCLUSIVE at this sample size. {self.band.null.capitalize()} "
f"would often score this well on this many rows, so this dataset "
"cannot tell the two apart. This is not a result in either "
"direction. Collect more rows to make the question answerable."
)
)
name = METRIC_NAMES.get(self.metric, self.metric.upper())
Expand Down
18 changes: 15 additions & 3 deletions src/plumbline/metrics/calibration.py
Original file line number Diff line number Diff line change
Expand Up @@ -597,10 +597,22 @@ def verdict(measured: float, floor: FloorBand) -> str:


def _judgment(measured: float, floor: FloorBand) -> str:
"""The trailing clause of a verdict, so one wording serves every caller."""
"""The trailing clause of a verdict, so one wording serves every caller.

The inconclusive wording puts the failure on the measurement, not on the
model. "Not distinguishable from a perfectly calibrated model" is literally
what the arithmetic says, and it reads as a pass: a reader skimming a
report sees their model compared to a perfect one and no difference found.
What it actually means is that this dataset is too small to resolve the
question either way, which is the absence of a result rather than a good
one. The sentence has to say so, because the sentence is what gets read.
"""
if is_distinguishable(measured, floor):
return "miscalibration is distinguishable from sampling noise."
return (
"not distinguishable from a perfectly calibrated model at this sample size. "
"Collect more rows before reading anything into it."
"INCONCLUSIVE at this sample size. A perfectly calibrated model would "
"often score this badly on this many rows, so this dataset cannot tell "
"the two apart. This is not a clean bill of health: nothing was "
"established either way. Collect more rows to make the question "
"answerable."
)
8 changes: 5 additions & 3 deletions src/plumbline/report/markdown.py
Original file line number Diff line number Diff line change
Expand Up @@ -148,9 +148,11 @@ def _how_to_read() -> list[str]:
"",
"- Every figure states the rows it was computed on and the null it is read "
"against. A number on its own is not a finding.",
'- "Not distinguishable" means the value sits inside what the null produces at '
"this sample size. It does not mean the systems are the same; it means this "
"dataset cannot tell them apart yet.",
'- **"INCONCLUSIVE" is not a pass.** It means the value sits inside what the '
"null already produces at this sample size, so this dataset cannot tell the two "
"apart. Nothing was established in either direction. A model that is genuinely "
"well calibrated and one that is badly calibrated can both land here on too few "
"rows, and the figure does not say which you have.",
"- Arms are grouped by what kind of number they report. **Figures in different "
"groups are not comparable** and are never placed side by side.",
'- "Not reported" is a result with a reason attached, not a missing cell.',
Expand Down
20 changes: 19 additions & 1 deletion tests/test_baseline.py
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,25 @@ def test_an_accuracy_at_chance_is_not_dressed_up_as_a_result() -> None:
figure = baseline.accuracy_figure([True] * 25 + [False] * 75, [4] * 100, n_boot=N_BOOT)

assert not figure.is_distinguishable
assert "not distinguishable" in figure.statement()
statement = figure.statement()
assert "INCONCLUSIVE" in statement
assert "cannot tell the two apart" in statement


def test_an_inconclusive_statement_does_not_read_as_a_pass() -> None:
"""The wording must not let a reader take "no difference found" as good news.

The old wording was "not distinguishable from <null> at this sample size",
which is what the arithmetic says and the opposite of what it means. A
reader skimming a report saw their model compared against a null and no
difference found, and read it as a clean result.
"""
figure = baseline.accuracy_figure([True] * 25 + [False] * 75, [4] * 100, n_boot=N_BOOT)
statement = figure.statement().lower()

assert "not a result in either direction" in statement
for pass_like in ("passes", "acceptable", "looks fine", "no issue", "is calibrated"):
assert pass_like not in statement


def test_auroc_of_an_uninformative_score_is_not_distinguishable_from_the_null() -> None:
Expand Down
6 changes: 5 additions & 1 deletion tests/test_calibration.py
Original file line number Diff line number Diff line change
Expand Up @@ -98,7 +98,11 @@ def test_a_calibrated_mock_lands_inside_its_own_noise_floor(n_cases: int) -> Non
measured_ece = calibration.ece(run.probabilities, run.outcomes)

assert not calibration.is_distinguishable(measured_ece, floors["ece"])
assert "not distinguishable" in calibration.verdict(measured_ece, floors["ece"])
verdict = calibration.verdict(measured_ece, floors["ece"])
assert "INCONCLUSIVE" in verdict
# The valence has to be explicit. Landing on the floor is the absence of a
# result, not a passing grade, and the sentence is what gets read.
assert "not a clean bill of health" in verdict


@pytest.mark.parametrize("n_cases", [500, 4000])
Expand Down
10 changes: 9 additions & 1 deletion tests/test_report.py
Original file line number Diff line number Diff line change
Expand Up @@ -145,7 +145,15 @@ def test_calibration_is_read_against_its_floor() -> None:

line = next(line for line in text.splitlines() if "ECE" in line)
assert "floor" in line
assert "distinguishable" in line
assert "INCONCLUSIVE" in line or "distinguishable from sampling noise" in line


def test_the_preamble_says_inconclusive_is_not_a_pass() -> None:
"""The report defines its own terms, because the report is what gets read."""
text = render(a_run())

assert '**"INCONCLUSIVE" is not a pass.**' in text
assert "Nothing was established in either direction." in text


def test_the_maximum_error_is_a_diagnostic_and_not_a_headline() -> None:
Expand Down
Loading