From fad7a119a8d6b7feb6371a9d28b2b344f865277b Mon Sep 17 00:00:00 2001 From: TMHSDigital <154358121+TMHSDigital@users.noreply.github.com> Date: Mon, 21 Sep 2026 22:33:50 -0400 Subject: [PATCH] docs: add a changelog and a pull request template CHANGELOG.md is seeded from the v0.1.0 release notes and carries an Unreleased section with what has landed since. Its preamble states the one thing a changelog for this project has to state: anything that changes what a number means is called out under Changed rather than buried under Fixed, because a figure that moved for a methodology reason is a different event from one that moved because it was wrong. The pull request template is four lines and adds no checklist. It asks for the why, and for the two confirmations this project actually cares about: that a rendered figure carries its n and its null, and that a new vendor was config rather than a new module. It also says which checks gate a merge, so a contributor seeing an advisory failure knows what to do with it. PLAN.md records what was considered and declined, with the reason and the condition that would reopen it: .editorconfig, pre-commit, CODEOWNERS, CITATION.cff and PyPI publishing. CI caching was not added because setup-uv already runs with enable-cache. A file added because projects usually have one is a file that rots, and a decision with no recorded reason gets re-proposed as an oversight. Co-Authored-By: Claude Opus 5 (1M context) --- .github/pull_request_template.md | 7 +++ CHANGELOG.md | 89 ++++++++++++++++++++++++++++++++ docs/PLAN.md | 32 ++++++++++++ 3 files changed, 128 insertions(+) create mode 100644 .github/pull_request_template.md create mode 100644 CHANGELOG.md diff --git a/.github/pull_request_template.md b/.github/pull_request_template.md new file mode 100644 index 0000000..9caea3a --- /dev/null +++ b/.github/pull_request_template.md @@ -0,0 +1,7 @@ +What this changes, and why. The diff says what; say why. + +If it renders a figure, confirm it carries its n and its null. If it adds a +vendor, confirm that was config rather than a new adapter module. + +CONTRIBUTING has the rest. CI gates the merge; CodeQL and Socket are advisory, +so say what you concluded if one of them fires. diff --git a/CHANGELOG.md b/CHANGELOG.md new file mode 100644 index 0000000..87ced80 --- /dev/null +++ b/CHANGELOG.md @@ -0,0 +1,89 @@ +# Changelog + +Notable changes per release. Dates are the release date, not the tag date when +those differ. + +This project is pre-1.0. The measurement behaviour is the stable part; the +Python API and the CLI flags are not, and a minor version may change either. +Anything that changes what a number **means** is called out under Changed, not +buried under Fixed, because a figure that moved for a methodology reason is a +different event from one that moved because it was wrong. + +## Unreleased + +### Fixed + +- The inconclusive verdict read as a pass. "Not distinguishable from a + perfectly calibrated model at this sample size" is what the arithmetic + establishes and close to the opposite of what it means, and readers took it + as a clean result. Every such figure now leads with `INCONCLUSIVE` and states + that nothing was established in either direction. **This changes report text, + not any number.** +- The package shipped no PEP 561 `py.typed` marker, so downstream type checkers + ignored its annotations and consumers silently saw `Any`. +- `plumbline version` printed `0.1.0.dev0` from the v0.1.0 release, because the + test asserting the version used a substring match that `0.1.0.dev0` satisfies. +- `plumbline report` on a missing or malformed artifact raised a bare + `FileNotFoundError` traceback instead of refusing with a reason. +- An adapter built without a required setting raised a bare `TypeError` from + `__init__` instead of naming the setting. +- CI ran every pull request twice, because both the `push` and `pull_request` + triggers fired on a branch pushed to origin. + +### Documentation + +- README rewritten for a reader arriving from a link: what the tool is now + precedes what it is not, and decision model, calibration, cascade, Noul, + binning noise and the Jev wire format are each defined where they appear. + Adds prerequisites, a bash quickstart, and a note that this is not on PyPI. +- The adapters table gained a "Run for real" column. Two of the three + transports have never run outside the test suite. +- SECURITY.md distinguishes what secret scanning covers from what it does not, + since the closest thing to a disclosure this project has had was a vendor's + price, which no scanner recognises. + +## v0.1.0 — 2026-09-21 + +First release. + +### Added + +- **Calibration measured against its own null.** ECE, MCE and Brier are each + reported beside the floor a perfectly calibrated model would produce at the + same row count, computed by simulation. A figure inside its floor is reported + as unresolvable rather than as a result. +- **Choice and Noul question types.** A yes/no row is asked as a Noul where the + transport has one, and every record carries both what the row asks and how it + was asked, so the two are never averaged together. +- **Three adapter transports.** `typesafe_wire` for the Jev wire format over + HTTP, `local_logits` for option-token logits from a pinned local checkpoint, + and `generative` as a text-generating control arm. Plus a seeded `mock`. +- **Three probability semantics classes.** `calibrated_claim`, + `restricted_softmax` and `none`. The report groups on this field and refuses + to place figures from different classes side by side. +- **Temperature scaling with a refusal gate.** Fitted on a held-out split. + When the residual says temperature is the wrong correction, or the split has + fewer than 200 rows, the tool emits no temperature rather than one that does + not fit. +- **Cascade threshold selection.** Given the cost of one escalation and one + wrong answer, it states where to cut and what that buys. Without both + numbers it refuses, because no benchmark can know them. +- **Cost from reported tokens against a dated pricing table.** Every entry + carries its source and the date it was read. A blank cost column names which + of four reasons made it blank. +- **Operator-supplied pricing** via `--pricing`, for vendors whose terms treat + their rates as confidential. Those ship unpriced, naming the page to read. +- **A loader that refuses rather than repairs.** A row whose gold label is not + among its own options is refused with its line number, because scoring it + would mark every system wrong and read as a model failure. +- **Latency percentiles by nearest rank**, and a results artifact recording the + requested model, the model that answered, the dataset hash and row count, and + the pricing entry applied, with credentials redacted. +- Apache-2.0. The vendored JevBench fixture is MIT and attributed in + `datasets/public/README.md`. + +### Known limitations at release + +Ordinal Score rows load but are excluded from every figure. One request per +case, so cost and latency are conservative relative to batched use. Temperature +scaling only. Verified on Windows and Ubuntu, Python 3.12 and 3.13. diff --git a/docs/PLAN.md b/docs/PLAN.md index 417ee77..1197669 100644 --- a/docs/PLAN.md +++ b/docs/PLAN.md @@ -125,6 +125,38 @@ Settled during the build. Reopen one only with a reason, not from scratch. control for it is the `--pricing` design plus a test. SECURITY.md says so in full, because a green scanning badge invites the wrong assumption. +### Repository files and tooling deliberately not added + +Considered and declined on 2026-09-21. Listed so they are not re-proposed as +oversights. Each would be defensible later for a stated reason; none is +defensible merely because projects usually have one. + +- **`.editorconfig`.** ruff already owns formatting here, CI enforces + `ruff format --check`, and no tool in this repository reads an editorconfig. + Adding one creates a second source of truth for line length and indentation + that can silently disagree with the first. `.gitattributes` already pins the + vendored fixture's bytes, which is the only line-ending rule that affects + correctness. Revisit if a contributor's editor is actually fighting ruff. +- **pre-commit.** It would catch exactly what CI already catches, in exchange + for a setup step in CONTRIBUTING, a pinned-hook config to keep current, and a + second place where the lint versions live. `uv run ruff check . && uv run + mypy --strict` is already documented and is one command. Revisit if CI + minutes or review round-trips become the bottleneck, which at this size they + are not. +- **CODEOWNERS.** One maintainer. The file would assign every path to the + person who would be reviewing it anyway, and the ruleset already requires a + pull request. Revisit on the second maintainer. +- **`CITATION.cff`.** Nobody has cited this. A citation file asserting how to + cite work nobody has referenced is a claim about its significance rather than + a service to a reader. Revisit if someone references the METHODOLOGY results. +- **PyPI publishing.** Premature. It commits the project to a name and to a + release cadence before the API has settled, and the API is explicitly not + stable before v0.2. The README says cloning is the install path, which is + honest and costs a reader one command. Revisit when the CLI flags stop + moving. +- **CI dependency caching** was not added because it was already there: + `astral-sh/setup-uv` runs with `enable-cache: true`. + ## Live validation Run on **2026-09-21**. Model requested `jev-latest`; model that answered