From db8e9aaff19671f2b30c3b8e02e9a9c1a97b76a3 Mon Sep 17 00:00:00 2001 From: TMHSDigital <154358121+TMHSDigital@users.noreply.github.com> Date: Wed, 23 Sep 2026 12:51:13 -0400 Subject: [PATCH] docs: stop using -- as a dash, and fail CI on it -- is for command lines: flags, git's end-of-options separator, and the like. It had been standing in for a dash in 26 places: METHODOLOGY, the docstrings and comments of six source files, a test, an example, and one report string, the load summary's note on ordinal-score rows. Each now uses a comma, colon, or parentheses. The example report carries the same note, so it changes with it; the site build still finds it identical, line for line, to what its command prints. A second step in CI's prose job fails on a bare -- between words, or one opening a continuation line, across the markdown, src/, tests/, examples/, site/, and scripts/. A flag has no space after its dashes, and HTML comments and table rules have no word before them, so none of those trip it. It finds all 26 on main and nothing now. The site smoke test's failure marker was a "--" too; it is now FAIL. Closes #54. This changes report text, not any number. Co-Authored-By: Claude Opus 5.5 (1M context) --- .github/workflows/ci.yml | 13 +++++++++++++ CHANGELOG.md | 6 ++++++ METHODOLOGY.md | 22 +++++++++++----------- conftest.py | 4 ++-- docs/example-report.md | 2 +- examples/smoke_public_dataset.py | 4 ++-- scripts/smoke_site.mjs | 4 ++-- src/plumbline/config.py | 2 +- src/plumbline/datasets/loader.py | 8 ++++---- src/plumbline/metrics/baseline.py | 10 +++++----- src/plumbline/report/markdown.py | 2 +- src/plumbline/types.py | 2 +- tests/test_local_logits.py | 2 +- 13 files changed, 50 insertions(+), 31 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 565feff..dd6fc47 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -149,6 +149,19 @@ jobs: exit 1 fi + # -- is for command lines (flags, and git's end-of-options), never a stand-in + # for a dash in prose. This finds a bare -- with words on both sides, or one + # opening a continuation line; a flag like --adapter has no space after it, + # and HTML comments and table rules have no word before it. + - name: Prose uses real punctuation, not -- as a dash + run: | + pattern='[\w,.;:)\]]\s+--\s+[A-Za-z(]|^\s*--\s+[a-z]' + if git grep -n -P "$pattern" -- '*.md' 'src/' 'tests/' 'examples/' 'conftest.py' \ + 'site/' 'scripts/' ':!datasets/public/' ':!site/vendor/'; then + echo "::error::-- used as a dash above; use a comma, colon, parentheses, or a new sentence" + exit 1 + fi + # Gate 6 found that the built distribution is a different artifact from the # source tree: the entry point, the optional extra boundary, and py.typed all # only exist once packaged. diff --git a/CHANGELOG.md b/CHANGELOG.md index bcc9e24..83db89f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -74,6 +74,12 @@ different event from one that moved because it was wrong. changes report text, not any number.** The job also covers `site/`, `scripts/`, and `.github/`, and the HTML entity and JavaScript escape spellings of the character. +- `--` is no longer used as a stand-in for a dash: METHODOLOGY, the source + comments and docstrings, and one report note (the ordinal-score line in the + load summary) now use commas, colons, or parentheses. A second check in the + same CI job fails on a bare `--` between words; flags, git's end-of-options + separator, and HTML comments are unaffected. **This changes report text, not + any number.** ### Fixed diff --git a/METHODOLOGY.md b/METHODOLOGY.md index 30af29f..3b7bb43 100644 --- a/METHODOLOGY.md +++ b/METHODOLOGY.md @@ -65,8 +65,8 @@ answer meets that condition by construction. That is no longer hypothetical. plumbline now asks a yes/no row as a Noul, so every such row is measured as one probability with nothing behind it. Those rows -are outside the multiclass Brier column -- the column renders as not reported for -them, never as zero -- and the temperature that can be fitted for them is the +are outside the multiclass Brier column (the column renders as not reported for +them, never as zero), and the temperature that can be fitted for them is the one-parameter approximation in the table above, the form that left ECE at five times the floor on the underconfident case. The penalty is a property of the answer shape, not of the model that produced it, and it is the price of a wire @@ -126,7 +126,7 @@ guard cannot bound a run it cannot cost. A Noul returns one number: the probability that the answer is yes. There is no distribution behind it and therefore no confidence statistic computed from one. -That makes it the cleanest calibration target in the API -- nothing is +That makes it the cleanest calibration target in the API: nothing is renormalized, nothing is derived, and `prob_selected` is exactly what the vendor reported. @@ -156,7 +156,7 @@ not exist for an answer with no distribution. ## Ordinal score questions are not scored in v0.1 Some datasets ask for a level rather than a label: 0, 1, 2, or 3 daily-rest -violations. The levels are ordered, and every metric here is rank-blind -- being +violations. The levels are ordered, and every metric here is rank-blind: being wrong by one level and wrong by three score identically. Flattening the levels into unordered options would discard exactly the structure that makes the question a score, so plumbline loads those rows, marks them, and leaves them out @@ -191,9 +191,9 @@ own mix of option widths, which is not 1/n for any single n once the widths differ. AUROC is read against a permutation null that keeps the observed ties and class balance. -Nothing prints as a bare number. A figure whose null cannot be built -- MCE when -no bin holds enough rows, AUROC when every case is correct or every case is wrong --- is reported as not reported with the reason, rather than as a number standing +Nothing prints as a bare number. A figure whose null cannot be built (MCE when +no bin holds enough rows, AUROC when every case is correct or every case is wrong) +is reported as not reported with the reason, rather than as a number standing on its own. The row count travels with the figure in the types as well as on the page: @@ -356,7 +356,7 @@ not comparable. A temperature is fitted on one half of the rows and reported on the other. The split is disjoint, the disjointness is checked rather than trusted, and the seed -and both sizes print in the report -- including when the verdict is a refusal, +and both sizes print in the report, including when the verdict is a refusal, because the procedure is part of the result. Fitting and reporting on the same rows manufactures an improvement that does not @@ -394,12 +394,12 @@ model of the same sharpness, which is a different and more useful claim than Every case is one request. Nothing is sampled repeatedly and voted on, nothing is retried to get a better-looking answer, and nothing is asked twice to reduce -variance -- because a number produced that way is not the number the system would +variance, because a number produced that way is not the number the system would give in production, and it is the production number plumbline is measuring. Transport failures are retried, with backoff, because a connection reset is not -an answer. Decisions are not. A refusal -- an option that is not a single token, -a generator that answered off-label, a model that declined -- is recorded once +an answer. Decisions are not. A refusal (an option that is not a single token, +a generator that answered off-label, a model that declined) is recorded once and never retried, since retrying would understate exactly the failure rate the report is there to show. A cache hit replaces the call entirely and is recorded as a hit, contributing to neither cost nor latency, because it measures disk. diff --git a/conftest.py b/conftest.py index e0263f3..772d3c6 100644 --- a/conftest.py +++ b/conftest.py @@ -3,8 +3,8 @@ ``tmp_path`` is only as safe as the temp root pytest resolves. Some sandboxed runners hand Python the working directory as that root, so ``tempfile.gettempdir()`` returns the repository itself and pytest builds -``pytest-of-/`` inside the working tree. The tests are not at fault -- -they all take ``tmp_path`` and never name a relative path -- so the fix belongs +``pytest-of-/`` inside the working tree. The tests are not at fault: +they all take ``tmp_path`` and never name a relative path, so the fix belongs here: pin the temp root once, before any fixture reads it, and refuse to run rather than scatter scratch files through the repo. diff --git a/docs/example-report.md b/docs/example-report.md index 8c521c7..af3590c 100644 --- a/docs/example-report.md +++ b/docs/example-report.md @@ -59,7 +59,7 @@ Generated 2026-09-22 against dataset `c18e9496`, 105 rows. 1 arm(s). ## Dataset -- 111 rows read from datasets\public\jevbench-hard.jsonl, 111 loaded, 0 refused. Translated from JevBench: the case text is the row's question above its state, and every row is asked as a one-of-n choice. plumbline's harness, prompts and scoring differ from JevBench's, so these numbers are not comparable with theirs. 6 rows wrote the gold label as a JSON number against string options; each was matched to the option of the same name. 38 rows carried criteria that do not describe the options one for one, so their option descriptions were dropped rather than guessed. 6 rows ask for an ordinal score. plumbline v0.1 has no ordinal support -- flattening levels into unordered options discards the ordering -- so they are loaded, marked, and excluded from scored results. +- 111 rows read from datasets\public\jevbench-hard.jsonl, 111 loaded, 0 refused. Translated from JevBench: the case text is the row's question above its state, and every row is asked as a one-of-n choice. plumbline's harness, prompts and scoring differ from JevBench's, so these numbers are not comparable with theirs. 6 rows wrote the gold label as a JSON number against string options; each was matched to the option of the same name. 38 rows carried criteria that do not describe the options one for one, so their option descriptions were dropped rather than guessed. 6 rows ask for an ordinal score. plumbline v0.1 has no ordinal support: flattening levels into unordered options discards the ordering, so they are loaded, marked, and excluded from scored results. - 6 score rows are excluded from every figure below: plumbline v0.1 scores choice and yes/no questions only. diff --git a/examples/smoke_public_dataset.py b/examples/smoke_public_dataset.py index cf0ff36..72f2680 100644 --- a/examples/smoke_public_dataset.py +++ b/examples/smoke_public_dataset.py @@ -6,9 +6,9 @@ wrote. It makes no network call and spends nothing, and the numbers it prints are meaningless: a seeded mock is answering, so its accuracy and its ECE are properties of the mock's configuration and nothing else. What it proves is that -the pieces compose on real shapes -- structured states, options that are yes/no +the pieces compose on real shapes (structured states, options that are yes/no on some rows and five-way on others, gold labels written as numbers, ordinal -rows that v0.1 will not score -- before a live run turns mistakes into money. +rows that v0.1 will not score) before a live run turns mistakes into money. The same run is available as ``plumbline run --format jevbench``. This file stays because it is the shortest readable path through the library, and because diff --git a/scripts/smoke_site.mjs b/scripts/smoke_site.mjs index 7df3d6d..f5112a1 100644 --- a/scripts/smoke_site.mjs +++ b/scripts/smoke_site.mjs @@ -227,10 +227,10 @@ const failures = []; async function check(name, body) { try { await body(); - console.log(` ok ${name}`); + console.log(` ok ${name}`); } catch (error) { failures.push(`${name}: ${error.message}`); - console.log(` -- ${name}: ${error.message}`); + console.log(` FAIL ${name}: ${error.message}`); } } const expect = (condition, what) => { diff --git a/src/plumbline/config.py b/src/plumbline/config.py index dabb3b1..8192ea2 100644 --- a/src/plumbline/config.py +++ b/src/plumbline/config.py @@ -47,7 +47,7 @@ class PricingConfigError(PlumblineError): #: #: The vendor's customer agreement makes its pricing information confidential and #: overrides the usual public-knowledge carve-out, so plumbline does not restate -#: the numbers in a file it publishes -- not in a price field, not in a quotation, +#: the numbers in a file it publishes: not in a price field, not in a quotation, #: and not in a worked example whose token counts would let a reader divide one #: out. The page is public and the reader can go and read it. Supplying the #: figures is the operator's act, not plumbline's. diff --git a/src/plumbline/datasets/loader.py b/src/plumbline/datasets/loader.py index 4349fb6..3d3ddcc 100644 --- a/src/plumbline/datasets/loader.py +++ b/src/plumbline/datasets/loader.py @@ -15,7 +15,7 @@ Two readers live here. ``load_jsonl`` reads plumbline's own record shape, which is what ``datasets/private/`` holds. ``load_jevbench`` reads the JevBench public file, which is a fixture for proving the pipeline composes, not a reproduction -of anyone's benchmark -- see ``datasets/public/README.md``. +of anyone's benchmark (see ``datasets/public/README.md``). """ from __future__ import annotations @@ -71,8 +71,8 @@ def row_count(self) -> int: def scoreable(self) -> tuple[Case, ...]: """The cases v0.1 is willing to turn into numbers. - An ordinal score row is loaded and kept in ``cases`` -- nothing is lost - quietly -- and left out of here, because flattening its levels into + An ordinal score row is loaded and kept in ``cases`` (nothing is lost + quietly) and left out of here, because flattening its levels into unordered options discards the ordering that makes it a score. A run takes this set; the count that was held back is in the notes. """ @@ -223,7 +223,7 @@ def load_jevbench(path: Path | str) -> LoadReport: if unsupported: notes.append( f"{unsupported} rows ask for an ordinal score. plumbline v0.1 has no ordinal " - "support -- flattening levels into unordered options discards the ordering -- " + "support: flattening levels into unordered options discards the ordering, " "so they are loaded, marked, and excluded from scored results." ) diff --git a/src/plumbline/metrics/baseline.py b/src/plumbline/metrics/baseline.py index 1ecf482..39464b3 100644 --- a/src/plumbline/metrics/baseline.py +++ b/src/plumbline/metrics/baseline.py @@ -2,8 +2,8 @@ Accuracy and AUROC are read the same way ECE is: against what the number would be if nothing were happening, at this exact sample size. A bare 0.72 accuracy is -unreadable -- it is excellent on eight-way options and it is nothing on two-way -options -- and a bare AUROC of 0.58 over 100 rows is well inside what an +unreadable (it is excellent on eight-way options and it is nothing on two-way +options), and a bare AUROC of 0.58 over 100 rows is well inside what an uninformative score column produces by chance. Both nulls are simulated rather than assumed. Accuracy's null draws an answer @@ -112,8 +112,8 @@ def chance_band( return NullBand( metric="accuracy", null="chance", - # The expectation is known exactly -- it is the mean of 1/n over the - # cases -- so it is computed rather than estimated. Only the spread, + # The expectation is known exactly (it is the mean of 1/n over the + # cases), so it is computed rather than estimated. Only the spread, # which is what the sample size controls, needs the simulation. mean=float(probabilities.mean()), p95=float(np.percentile(draws, 95)), @@ -152,7 +152,7 @@ def auroc_band( """What AUROC does when the score column says nothing about the outcome. Built by permuting the outcomes against the same scores, so the band keeps - the observed ties and the observed class balance -- both of which move it. + the observed ties and the observed class balance, both of which move it. """ if len(scores) != len(correct): raise ValueError(f"{len(scores)} scores against {len(correct)} outcomes; these must match") diff --git a/src/plumbline/report/markdown.py b/src/plumbline/report/markdown.py index eb438a2..bb8a477 100644 --- a/src/plumbline/report/markdown.py +++ b/src/plumbline/report/markdown.py @@ -396,7 +396,7 @@ def _cascade_rows( A threshold set against a raw overconfident probability sits in the wrong place, because 0.9 from an overconfident model is not 0.9. So when a temperature was recommended, the threshold is chosen on the held-out rows - with that temperature applied -- the same rows the temperature was judged + with that temperature applied: the same rows the temperature was judged on, never the rows it was fitted on. """ predictions = [record.prediction for record in successes if record.prediction] diff --git a/src/plumbline/types.py b/src/plumbline/types.py index f78ff34..57deae4 100644 --- a/src/plumbline/types.py +++ b/src/plumbline/types.py @@ -138,7 +138,7 @@ class InsufficientDataError(PlumblineError): class DatasetError(PlumblineError): """A dataset could not be read, or a row in it could not be scored. - Raised for the file as a whole -- missing, empty, or demanded complete when + Raised for the file as a whole: missing, empty, or demanded complete when it is not. A single bad row is a refusal recorded in the load report rather than an exception, so one typo does not stop the other 110 rows from running. """ diff --git a/tests/test_local_logits.py b/tests/test_local_logits.py index 6d2ae79..4856302 100644 --- a/tests/test_local_logits.py +++ b/tests/test_local_logits.py @@ -2,7 +2,7 @@ This adapter is the null hypothesis of the whole tool: a restricted softmax over the option tokens, with no calibration claim attached to it. Everything here is -about keeping that claim honest -- the probabilities are a softmax over the +about keeping that claim honest: the probabilities are a softmax over the option tokens and nothing else, the checkpoint is pinned and recorded, and a case whose options do not map to single tokens is refused rather than quietly turned into a different question.