Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
240 changes: 182 additions & 58 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,12 +5,42 @@
Measure whether a decision model's probabilities are trustworthy on your own
labeled data, and decide what to do about it.

> **v0.1.0, one maintainer.** The measurement behaviour is settled; the Python
> API and the CLI flags are not, and will change in v0.2. Pin a version if you
> build on it.

## What it is

A **decision model** is one you ask a closed question and get back a label plus
a number: "which of these five categories is this ticket" together with "0.83".
That number is the point. If it is honest, your code can act on 0.95 and route
0.6 to a human, and you have a system. If it is not, you have a confident
guess in a trench coat.

plumbline is a command line tool that takes **your** labeled rows, runs one or
more models over them, and tells you three things:

1. **Is the number honest?** A model is *calibrated* when the things it calls
70% likely happen about 70% of the time. plumbline measures that, and
crucially measures it against what the same figure would look like if the
model were perfect, because on a few hundred rows those are closer than
anyone expects.
2. **Where do I set my threshold?** Act automatically above it, escalate below
it.
3. **What does that save me?** A **cascade** runs a cheap model first and sends
only the uncertain cases to an expensive one. Given what an escalation costs
you and what a mistake costs you, plumbline says where to cut and what you
get for it.

It is a measuring instrument. It has no opinion about which model you should
pick, and it will refuse to answer a question your data cannot support.

## What this is not

plumbline is not a leaderboard. It does not rank vendors, it publishes no
combined score, and it will not tell you which model is best.

If you want a cross-vendor ranking of Jev-class decision models, go to
If you want a cross-vendor ranking of decision models, go to
[JevBench](https://github.com/fstandhartinger/jevbench) and
[Benchmark Heaven](https://benchmarkheaven.com/jev-models). That is their job and
they do it properly, across many models, on a shared dataset, with a published
Expand All @@ -32,40 +62,57 @@ them.

## The argument

Expected Calibration Error has a floor that is not zero, and the floor depends on
how many rows you have. A perfectly calibrated model, measured on a few hundred
rows, does not score 0. It scores some positive number determined by binning
noise and sample size. If you do not know that number, you cannot read your own.
The standard way to score calibration is **Expected Calibration Error**: sort
the predictions into bins by confidence, and in each bin compare the claimed
confidence against how often the model was actually right. Average the gaps.
Zero would be perfect.

This is not a rounding concern. On a few hundred rows a calibration claim is
frequently not measurable at all.
Zero is not achievable, and that is the problem. With a finite number of rows,
each bin holds a handful of cases, and a handful of coin flips does not land
exactly on its own probability. That scatter, **binning noise**, puts a floor
under ECE that has nothing to do with the model. The floor rises as your row
count falls. **A perfectly calibrated model on a few hundred rows does not
score 0, and if you do not know what it would score, you cannot read your
own number.**

So plumbline reports every inferential figure against its own null, and says
plainly when a value sits inside what that null already produces.
This is not a rounding concern. On a few hundred rows, a calibration claim is
frequently not measurable at all.

**When it says a figure is inconclusive, that is not a pass.** It means the
dataset cannot tell your model apart from the null, so nothing was established
in either direction. A model that is genuinely well calibrated and one that is
badly calibrated can both land there on too few rows, and the figure does not
say which you have. Reading it as a clean bill of health is the single easiest
mistake to make with this tool, and it inverts the conclusion.
So plumbline computes that floor by simulation and prints every inferential
figure against it. Here is a real line from the example report, unedited:

From the example report, 105 rows:
> ECE 0.0740 over 105 rows (10 equal width bins), against a calibrated-model
> floor of 0.0707 (95th percentile 0.1109): INCONCLUSIVE at this sample size. A
> perfectly calibrated model would often score this badly on this many rows, so
> this dataset cannot tell the two apart. This is not a clean bill of health:
> nothing was established either way. Collect more rows to make the question
> answerable.

- ECE 0.0740, against a calibrated-model floor of 0.0707 and a 95th percentile of
0.1109. Inconclusive.
- Brier 0.1711, against a floor of 0.1489. Inconclusive.
- Confidence AUROC 0.6034, against a permutation null of 0.4997 and a 95th
percentile of 0.6116. Inconclusive.
- Accuracy 0.7714, against a chance null of 0.3416. Better than chance.
**Read "inconclusive" as an absence of a result, not a pass.** It means your
rows cannot tell your model apart from a perfect one, so nothing was
established in either direction. A genuinely well calibrated model and a badly
calibrated one both land there on too few rows, and the figure does not say
which you have. Taking it as a clean bill of health inverts the conclusion, and
it is the easiest mistake to make with this tool.

Four figures, one of which supports a conclusion. A tool that printed the first
three on their own would be handing you numbers that look like findings and are
not. That refusal is the product.
On that same 105-row run, three of the four headline figures came back
inconclusive and only accuracy cleared its null. A tool that printed the other
three alone would be handing you numbers that look like findings and are not.
That refusal is the product.

## Quickstart

PowerShell. Nothing below needs an API key or spends anything.
**Prerequisites:** Python 3.12 or later, and [uv](https://docs.astral.sh/uv/).
An older Python gives a resolver error rather than a clear message, so check
with `python --version` first.

**plumbline is not on PyPI.** `pip install plumbline` will not work. Clone the
repository; that is the intended install path for v0.1.

Nothing in this first section needs an API key or spends anything.

<details open>
<summary><b>PowerShell</b></summary>

```powershell
git clone https://github.com/TMHSDigital/plumbline
Expand All @@ -77,6 +124,23 @@ uv run plumbline run datasets/public/jevbench-hard.jsonl `
--results results --report results/report.md
```

</details>

<details>
<summary><b>bash or zsh</b></summary>

```bash
git clone https://github.com/TMHSDigital/plumbline
cd plumbline
uv sync

uv run plumbline run datasets/public/jevbench-hard.jsonl \
--adapter mock --format jevbench \
--results results --report results/report.md
```

</details>

That loads the vendored public fixture, runs a deterministic seeded mock over it,
computes every metric against its null, and writes both a results artifact and a
report. No network call. Both land in `results/`, which is gitignored, so
Expand All @@ -85,14 +149,44 @@ following this leaves your clone clean.
Expected output shape:

```
111 rows read from datasets\public\jevbench-hard.jsonl, 111 loaded, 0 refused. ...
artifact: results\20260921T222053+0000-mock-c18e9496.json
report: results\report.md
111 rows read from datasets/public/jevbench-hard.jsonl, 111 loaded, 0 refused. ...
artifact: results/20260921T222053+0000-mock-c18e9496.json
report: results/report.md
```

```
uv run plumbline adapters # what this install can run
uv run plumbline version
uv run plumbline run --help
```

### Running a local model, still without a key

`local_logits` reads option-token probabilities out of a checkpoint on your own
machine, so it needs no API key. It does need the optional `local` extra, which
a plain `uv sync` does not install:

```
uv sync --extra local
```

Without it every case fails with a message telling you this, so if a local run
reports no figures at all, that is the first thing to check.

### Running a hosted vendor

This one spends money. Set a key, name an adapter, and cap the run.

Cost needs a pricing table **you** supply, because plumbline ships no figures
for vendors whose terms treat pricing as confidential. Copy the template and
fill in the rates from the vendor's own page:

```
cp docs/pricing.example.json my-pricing.json
```

To run a real vendor instead, set a key and name an adapter. Cost needs a pricing
table you supply; see [Limitations](#limitations) and
[docs/pricing.example.json](docs/pricing.example.json).
<details open>
<summary><b>PowerShell</b></summary>

```powershell
$env:TYPESAFE_API_KEY = "your-key-here"
Expand All @@ -103,51 +197,78 @@ uv run plumbline run datasets/public/jevbench-hard.jsonl `
--max-cases 40 --pricing my-pricing.json
```

Your own data goes in `datasets/private/`, which is gitignored, and that is the
only path on which the recalibration numbers mean anything.
</details>

```powershell
uv run plumbline adapters # what this install can run
uv run plumbline version
<details>
<summary><b>bash or zsh</b></summary>

```bash
export TYPESAFE_API_KEY="your-key-here"

uv run plumbline run datasets/public/jevbench-hard.jsonl \
--adapter typesafe_wire --model jev-latest --format jevbench \
--results results --report results/report.md \
--max-cases 40 --pricing my-pricing.json
```

## Example report
</details>

[docs/example-report.md](docs/example-report.md) is real, unedited output. Four
lines from it:
Without a pricing table the run still works; cost reports as unpriced, and
`--max-cost-usd` refuses rather than bounding a run it cannot cost. See
[Limitations](#limitations).

> ECE 0.0740 over 105 rows (10 equal width bins), against a calibrated-model
> floor of 0.0707 (95th percentile 0.1109): INCONCLUSIVE at this sample size. A
> perfectly calibrated model would often score this badly on this many rows, so
> this dataset cannot tell the two apart. This is not a clean bill of health:
> nothing was established either way. Collect more rows to make the question
> answerable.
Your own data goes in `datasets/private/`, which is gitignored, and that is the
only path on which the recalibration numbers mean anything.

## Example report

> No cost available. None of the 105 cases could be priced, so cost is not
> reported rather than being shown as zero. 105 rows: tokens were reported, but
> the model that answered is not priced.
[docs/example-report.md](docs/example-report.md) is real, unedited output. The
ECE line is quoted in [The argument](#the-argument) above. Three more, each
showing the tool declining to do something:

> Not reported. recalibration needs at least 200 held-out evaluation rows and
> this split has 53. Fitting a temperature on fewer rows produces a number whose
> uncertainty is larger than the correction it claims to make, and it arrives
> looking like a measurement.

> Not reported. A threshold is decided by two numbers no benchmark can know:
> what one escalation to the expensive arm costs, and what one wrong answer
> costs.

> 38 noul rows were asked as choice questions, which is a different question from
> the one the dataset states. Not comparable with an arm that asked them as noul.

A **Noul** is a yes/no question that returns one probability directly, rather
than a distribution over options. Asking a yes/no row as a two-option choice is
a different question, so the report keeps the two apart instead of averaging
across the difference.

The arm in that report is the seeded mock, labeled as such at the top of the
page, so anyone can reproduce it with one command and no key.
page, so anyone can reproduce it with one command and no key. Its numbers are
properties of plumbline's harness, not a measurement of any vendor.

## Adapters and probability semantics

Three transports. Adding a vendor is config, not code.
Three real transports, plus a mock for smoke tests. **The intent is that adding
a vendor is config rather than code**, and that intent has so far been verified
against one endpoint; demonstrating it against a second is
[issue #11](https://github.com/TMHSDigital/plumbline/issues/11).

The "Run for real" column is deliberate. Two of these have only ever run
against test fakes, which is [issue #3](https://github.com/TMHSDigital/plumbline/issues/3)
and the highest-value item in the backlog.

The **Jev wire format** below is one vendor's HTTP shape for typed questions:
you send some state plus a question with declared options, and you get back a
selected option, a probability for each, and a confidence. Several open models
and self-hosted servers speak it, which is why one adapter covers all of them.

| Adapter | Transport | Semantics | Adding one |
|---|---|---|---|
| `typesafe_wire` | Jev wire format over HTTP | `calibrated_claim` | A `base_url`. Anything serving the Jev wire format is a config entry, including self-hosted endpoints and open models behind a compatible server. |
| `local_logits` | Option-token logits from a local checkpoint | `restricted_softmax` | A HuggingFace model id and a pinned revision. Requires the optional `local` extra. |
| `generative` | Chat completion, parsed | `none` | A model string. |
| `mock` | None, seeded | configurable | Built in. A deterministic stand-in for smoke tests. |
| Adapter | Transport | Semantics | Run for real | Adding one |
|---|---|---|---|---|
| `typesafe_wire` | Jev wire format over HTTP | `calibrated_claim` | Yes, 40 rows against a hosted vendor | A `base_url`. Anything serving the same wire format is a config entry, including self-hosted endpoints and open models behind a compatible server. |
| `local_logits` | Option-token logits from a local checkpoint | `restricted_softmax` | **No, tests only** | A HuggingFace model id and a pinned revision. Needs the optional `local` extra. |
| `generative` | Chat completion, parsed | `none` | **No, tests only** | A model string. |
| `mock` | None, seeded | configurable | Yes, it is the example report | Built in. A deterministic stand-in, not a system under test. |

The semantics classes are the whole reason the report refuses some comparisons:

Expand Down Expand Up @@ -195,7 +316,10 @@ Specific, and none of them are going to surprise you later.
- **Probabilities from a hosted API may arrive quantized.** That bounds the
resolution of any threshold or bin computed from them. METHODOLOGY says what
the bound is and where it bites.
- **Windows is the only verified platform** until CI says otherwise.
- **Verified on Windows and Ubuntu, Python 3.12 and 3.13.** macOS is untested.
- **Two of the three transports have never run outside the test suite.** See
the adapters table above and
[issue #3](https://github.com/TMHSDigital/plumbline/issues/3).

## Related work

Expand Down
Loading