Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -149,6 +149,19 @@ jobs:
exit 1
fi

# -- is for command lines (flags, and git's end-of-options), never a stand-in
# for a dash in prose. This finds a bare -- with words on both sides, or one
# opening a continuation line; a flag like --adapter has no space after it,
# and HTML comments and table rules have no word before it.
- name: Prose uses real punctuation, not -- as a dash
run: |
pattern='[\w,.;:)\]]\s+--\s+[A-Za-z(]|^\s*--\s+[a-z]'
if git grep -n -P "$pattern" -- '*.md' 'src/' 'tests/' 'examples/' 'conftest.py' \
'site/' 'scripts/' ':!datasets/public/' ':!site/vendor/'; then
echo "::error::-- used as a dash above; use a comma, colon, parentheses, or a new sentence"
exit 1
fi

# Gate 6 found that the built distribution is a different artifact from the
# source tree: the entry point, the optional extra boundary, and py.typed all
# only exist once packaged.
Expand Down
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,6 +74,12 @@ different event from one that moved because it was wrong.
changes report text, not any number.** The job also covers `site/`,
`scripts/`, and `.github/`, and the HTML entity and JavaScript escape
spellings of the character.
- `--` is no longer used as a stand-in for a dash: METHODOLOGY, the source
comments and docstrings, and one report note (the ordinal-score line in the
load summary) now use commas, colons, or parentheses. A second check in the
same CI job fails on a bare `--` between words; flags, git's end-of-options
separator, and HTML comments are unaffected. **This changes report text, not
any number.**

### Fixed

Expand Down
22 changes: 11 additions & 11 deletions METHODOLOGY.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,8 +65,8 @@ answer meets that condition by construction.

That is no longer hypothetical. plumbline now asks a yes/no row as a Noul, so
every such row is measured as one probability with nothing behind it. Those rows
are outside the multiclass Brier column -- the column renders as not reported for
them, never as zero -- and the temperature that can be fitted for them is the
are outside the multiclass Brier column (the column renders as not reported for
them, never as zero), and the temperature that can be fitted for them is the
one-parameter approximation in the table above, the form that left ECE at five
times the floor on the underconfident case. The penalty is a property of the
answer shape, not of the model that produced it, and it is the price of a wire
Expand Down Expand Up @@ -126,7 +126,7 @@ guard cannot bound a run it cannot cost.

A Noul returns one number: the probability that the answer is yes. There is no
distribution behind it and therefore no confidence statistic computed from one.
That makes it the cleanest calibration target in the API -- nothing is
That makes it the cleanest calibration target in the API: nothing is
renormalized, nothing is derived, and `prob_selected` is exactly what the vendor
reported.

Expand Down Expand Up @@ -156,7 +156,7 @@ not exist for an answer with no distribution.
## Ordinal score questions are not scored in v0.1

Some datasets ask for a level rather than a label: 0, 1, 2, or 3 daily-rest
violations. The levels are ordered, and every metric here is rank-blind -- being
violations. The levels are ordered, and every metric here is rank-blind: being
wrong by one level and wrong by three score identically. Flattening the levels
into unordered options would discard exactly the structure that makes the
question a score, so plumbline loads those rows, marks them, and leaves them out
Expand Down Expand Up @@ -191,9 +191,9 @@ own mix of option widths, which is not 1/n for any single n once the widths
differ. AUROC is read against a permutation null that keeps the observed ties and
class balance.

Nothing prints as a bare number. A figure whose null cannot be built -- MCE when
no bin holds enough rows, AUROC when every case is correct or every case is wrong
-- is reported as not reported with the reason, rather than as a number standing
Nothing prints as a bare number. A figure whose null cannot be built (MCE when
no bin holds enough rows, AUROC when every case is correct or every case is wrong)
is reported as not reported with the reason, rather than as a number standing
on its own.

The row count travels with the figure in the types as well as on the page:
Expand Down Expand Up @@ -356,7 +356,7 @@ not comparable.

A temperature is fitted on one half of the rows and reported on the other. The
split is disjoint, the disjointness is checked rather than trusted, and the seed
and both sizes print in the report -- including when the verdict is a refusal,
and both sizes print in the report, including when the verdict is a refusal,
because the procedure is part of the result.

Fitting and reporting on the same rows manufactures an improvement that does not
Expand Down Expand Up @@ -394,12 +394,12 @@ model of the same sharpness, which is a different and more useful claim than

Every case is one request. Nothing is sampled repeatedly and voted on, nothing is
retried to get a better-looking answer, and nothing is asked twice to reduce
variance -- because a number produced that way is not the number the system would
variance, because a number produced that way is not the number the system would
give in production, and it is the production number plumbline is measuring.

Transport failures are retried, with backoff, because a connection reset is not
an answer. Decisions are not. A refusal -- an option that is not a single token,
a generator that answered off-label, a model that declined -- is recorded once
an answer. Decisions are not. A refusal (an option that is not a single token,
a generator that answered off-label, a model that declined) is recorded once
and never retried, since retrying would understate exactly the failure rate the
report is there to show. A cache hit replaces the call entirely and is recorded
as a hit, contributing to neither cost nor latency, because it measures disk.
Expand Down
4 changes: 2 additions & 2 deletions conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,8 +3,8 @@
``tmp_path`` is only as safe as the temp root pytest resolves. Some sandboxed
runners hand Python the working directory as that root, so
``tempfile.gettempdir()`` returns the repository itself and pytest builds
``pytest-of-<user>/`` inside the working tree. The tests are not at fault --
they all take ``tmp_path`` and never name a relative path -- so the fix belongs
``pytest-of-<user>/`` inside the working tree. The tests are not at fault:
they all take ``tmp_path`` and never name a relative path, so the fix belongs
here: pin the temp root once, before any fixture reads it, and refuse to run
rather than scatter scratch files through the repo.

Expand Down
2 changes: 1 addition & 1 deletion docs/example-report.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,7 @@ Generated 2026-09-22 against dataset `c18e9496`, 105 rows. 1 arm(s).

## Dataset

- 111 rows read from datasets\public\jevbench-hard.jsonl, 111 loaded, 0 refused. Translated from JevBench: the case text is the row's question above its state, and every row is asked as a one-of-n choice. plumbline's harness, prompts and scoring differ from JevBench's, so these numbers are not comparable with theirs. 6 rows wrote the gold label as a JSON number against string options; each was matched to the option of the same name. 38 rows carried criteria that do not describe the options one for one, so their option descriptions were dropped rather than guessed. 6 rows ask for an ordinal score. plumbline v0.1 has no ordinal support -- flattening levels into unordered options discards the ordering -- so they are loaded, marked, and excluded from scored results.
- 111 rows read from datasets\public\jevbench-hard.jsonl, 111 loaded, 0 refused. Translated from JevBench: the case text is the row's question above its state, and every row is asked as a one-of-n choice. plumbline's harness, prompts and scoring differ from JevBench's, so these numbers are not comparable with theirs. 6 rows wrote the gold label as a JSON number against string options; each was matched to the option of the same name. 38 rows carried criteria that do not describe the options one for one, so their option descriptions were dropped rather than guessed. 6 rows ask for an ordinal score. plumbline v0.1 has no ordinal support: flattening levels into unordered options discards the ordering, so they are loaded, marked, and excluded from scored results.
- 6 score rows are excluded from every figure below: plumbline v0.1 scores choice and yes/no questions only.


Expand Down
4 changes: 2 additions & 2 deletions examples/smoke_public_dataset.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,9 +6,9 @@
wrote. It makes no network call and spends nothing, and the numbers it prints
are meaningless: a seeded mock is answering, so its accuracy and its ECE are
properties of the mock's configuration and nothing else. What it proves is that
the pieces compose on real shapes -- structured states, options that are yes/no
the pieces compose on real shapes (structured states, options that are yes/no
on some rows and five-way on others, gold labels written as numbers, ordinal
rows that v0.1 will not score -- before a live run turns mistakes into money.
rows that v0.1 will not score) before a live run turns mistakes into money.

The same run is available as ``plumbline run --format jevbench``. This file stays
because it is the shortest readable path through the library, and because
Expand Down
4 changes: 2 additions & 2 deletions scripts/smoke_site.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -227,10 +227,10 @@ const failures = [];
async function check(name, body) {
try {
await body();
console.log(` ok ${name}`);
console.log(` ok ${name}`);
} catch (error) {
failures.push(`${name}: ${error.message}`);
console.log(` -- ${name}: ${error.message}`);
console.log(` FAIL ${name}: ${error.message}`);
}
}
const expect = (condition, what) => {
Expand Down
2 changes: 1 addition & 1 deletion src/plumbline/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ class PricingConfigError(PlumblineError):
#:
#: The vendor's customer agreement makes its pricing information confidential and
#: overrides the usual public-knowledge carve-out, so plumbline does not restate
#: the numbers in a file it publishes -- not in a price field, not in a quotation,
#: the numbers in a file it publishes: not in a price field, not in a quotation,
#: and not in a worked example whose token counts would let a reader divide one
#: out. The page is public and the reader can go and read it. Supplying the
#: figures is the operator's act, not plumbline's.
Expand Down
8 changes: 4 additions & 4 deletions src/plumbline/datasets/loader.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@
Two readers live here. ``load_jsonl`` reads plumbline's own record shape, which
is what ``datasets/private/`` holds. ``load_jevbench`` reads the JevBench public
file, which is a fixture for proving the pipeline composes, not a reproduction
of anyone's benchmark -- see ``datasets/public/README.md``.
of anyone's benchmark (see ``datasets/public/README.md``).
"""

from __future__ import annotations
Expand Down Expand Up @@ -71,8 +71,8 @@ def row_count(self) -> int:
def scoreable(self) -> tuple[Case, ...]:
"""The cases v0.1 is willing to turn into numbers.

An ordinal score row is loaded and kept in ``cases`` -- nothing is lost
quietly -- and left out of here, because flattening its levels into
An ordinal score row is loaded and kept in ``cases`` (nothing is lost
quietly) and left out of here, because flattening its levels into
unordered options discards the ordering that makes it a score. A run
takes this set; the count that was held back is in the notes.
"""
Expand Down Expand Up @@ -223,7 +223,7 @@ def load_jevbench(path: Path | str) -> LoadReport:
if unsupported:
notes.append(
f"{unsupported} rows ask for an ordinal score. plumbline v0.1 has no ordinal "
"support -- flattening levels into unordered options discards the ordering -- "
"support: flattening levels into unordered options discards the ordering, "
"so they are loaded, marked, and excluded from scored results."
)

Expand Down
10 changes: 5 additions & 5 deletions src/plumbline/metrics/baseline.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,8 +2,8 @@

Accuracy and AUROC are read the same way ECE is: against what the number would
be if nothing were happening, at this exact sample size. A bare 0.72 accuracy is
unreadable -- it is excellent on eight-way options and it is nothing on two-way
options -- and a bare AUROC of 0.58 over 100 rows is well inside what an
unreadable (it is excellent on eight-way options and it is nothing on two-way
options), and a bare AUROC of 0.58 over 100 rows is well inside what an
uninformative score column produces by chance.

Both nulls are simulated rather than assumed. Accuracy's null draws an answer
Expand Down Expand Up @@ -112,8 +112,8 @@ def chance_band(
return NullBand(
metric="accuracy",
null="chance",
# The expectation is known exactly -- it is the mean of 1/n over the
# cases -- so it is computed rather than estimated. Only the spread,
# The expectation is known exactly (it is the mean of 1/n over the
# cases), so it is computed rather than estimated. Only the spread,
# which is what the sample size controls, needs the simulation.
mean=float(probabilities.mean()),
p95=float(np.percentile(draws, 95)),
Expand Down Expand Up @@ -152,7 +152,7 @@ def auroc_band(
"""What AUROC does when the score column says nothing about the outcome.

Built by permuting the outcomes against the same scores, so the band keeps
the observed ties and the observed class balance -- both of which move it.
the observed ties and the observed class balance, both of which move it.
"""
if len(scores) != len(correct):
raise ValueError(f"{len(scores)} scores against {len(correct)} outcomes; these must match")
Expand Down
2 changes: 1 addition & 1 deletion src/plumbline/report/markdown.py
Original file line number Diff line number Diff line change
Expand Up @@ -396,7 +396,7 @@ def _cascade_rows(
A threshold set against a raw overconfident probability sits in the wrong
place, because 0.9 from an overconfident model is not 0.9. So when a
temperature was recommended, the threshold is chosen on the held-out rows
with that temperature applied -- the same rows the temperature was judged
with that temperature applied: the same rows the temperature was judged
on, never the rows it was fitted on.
"""
predictions = [record.prediction for record in successes if record.prediction]
Expand Down
2 changes: 1 addition & 1 deletion src/plumbline/types.py
Original file line number Diff line number Diff line change
Expand Up @@ -138,7 +138,7 @@ class InsufficientDataError(PlumblineError):
class DatasetError(PlumblineError):
"""A dataset could not be read, or a row in it could not be scored.

Raised for the file as a whole -- missing, empty, or demanded complete when
Raised for the file as a whole: missing, empty, or demanded complete when
it is not. A single bad row is a refusal recorded in the load report rather
than an exception, so one typo does not stop the other 110 rows from running.
"""
Expand Down
2 changes: 1 addition & 1 deletion tests/test_local_logits.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

This adapter is the null hypothesis of the whole tool: a restricted softmax over
the option tokens, with no calibration claim attached to it. Everything here is
about keeping that claim honest -- the probabilities are a softmax over the
about keeping that claim honest: the probabilities are a softmax over the
option tokens and nothing else, the checkpoint is pinned and recorded, and a
case whose options do not map to single tokens is refused rather than quietly
turned into a different question.
Expand Down
Loading