Skip to content

docs(readme): rewrite for someone who arrived from a link - #14

Merged
TMHSDigital merged 1 commit into
mainfrom
docs/readme-for-a-stranger
Sep 22, 2026
Merged

TMHSDigital merged 1 commit into
mainfrom
docs/readme-for-a-stranger

Conversation

@TMHSDigital

Copy link
Copy Markdown
Owner

Everything from the Part 3 review, plus the six additions.

The review's central finding was that 195 words passed before the tool was
positively described
, and four terms appeared as if common knowledge: Jev,
JevBench, Benchmark Heaven, cascade. "Decision model", the class of thing
this tool measures, was never defined at all.

Structural

  • "What it is" now precedes "What this is not", which stays where it was.
    You cannot appreciate "not a leaderboard" before knowing what the thing is,
    and that framing was previously the entire first screen.
  • The argument leads with the real report line as its evidence, instead of
    paraphrasing the same figures as prose bullets and quoting them again sixty
    lines later. The figures appeared twice in two formats and the version with
    authority was the less prominent one.
  • bash/zsh quickstart alongside PowerShell, in collapsible sections. CI
    proves Ubuntu works, and most people who would use a Python evaluation tool
    are not on Windows.

Definitions added

decision model, calibration, cascade, Noul, binning noise, and the
Jev wire format. ECE is now explained as a procedure before its floor is
discussed, because the floor argument is unreadable if you do not know what is
being floored.

Overstatements corrected

From the independent review of the rendered page

  • --pricing my-pricing.json pointed at a file that never existed. Anyone
    copying it failed on their first paid run. It now shows
    cp docs/pricing.example.json my-pricing.json.
  • Prerequisites stated: Python 3.12+ and uv, with a note to check
    python --version, since 3.11 gives a resolver error rather than a message.
  • Says plainly it is not on PyPI and that cloning is the install path.
  • Maturity note near the top: v0.1.0, one maintainer, API will change in
    v0.2.
  • uv sync --extra local documented. Without it the local arm fails every
    case.
  • The excerpt no longer leads with the cost line. You were right that it
    read oddly: a seeded mock obviously has no price, so it demonstrated nothing.
    All three quotes now show the tool declining to answer.

Kept on your instruction

"Trustworthy" in the opening line, as a fair summary of calibration plus
discrimination. The GitHub description is aligned to match it.

Verified

All four block quotes checked whitespace-normalised against the committed
report: verbatim. Both internal anchors resolve. No em dashes, no emoji. Gate
green.

The README was written by someone who already knew why the project exists, and
it showed. 195 words passed before the tool was positively described, and four
terms appeared as if common knowledge: Jev, JevBench, Benchmark Heaven and
cascade. "Decision model", the class of thing this tool measures, was never
defined at all.

What changed.

A "What it is" section now precedes "What this is not", which stays where it
was. You cannot appreciate "not a leaderboard" before knowing what the thing
is, and the not-a-leaderboard framing was previously the entire first screen.
Decision model, calibration, cascade and Noul are each defined in a sentence
where they first appear, and the Jev wire format gets a gloss in the adapters
section. ECE is explained as a procedure before its floor is discussed, and
binning noise is named rather than assumed.

The argument now leads with the real report line as its evidence instead of
paraphrasing the same figures as prose bullets and quoting them again sixty
lines later. The figures appeared twice in two formats; the version with
authority was the less prominent one.

Quickstart states its prerequisites, says plainly that this is not on PyPI and
that cloning is the install path, and offers bash alongside PowerShell. CI
proves Ubuntu works and most people who would use a Python evaluation tool are
not on Windows. It also documents `uv sync --extra local`, without which the
local arm fails every case, and it no longer tells the reader to pass
`--pricing my-pricing.json`, a file that never existed. It shows how to make
one.

Three overstatements corrected. "Adding a vendor is config, not code" is
labeled as intent verified against one endpoint, linking #11. The adapters
table gains a "Run for real" column, because two of the three transports have
never left the test suite and presenting four rows as equally available was an
omission that functioned as a claim. "Windows is the only verified platform"
was stale: CI covers Ubuntu too.

Adds a maturity note near the top: v0.1.0, one maintainer, the API will change
in v0.2. The GitHub description is aligned with the opening line.

The example report excerpt no longer leads with the cost line, which read
oddly: a seeded mock obviously has no price, so it demonstrated nothing. The
three quotes now all show the tool declining to answer, which is the point.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
@TMHSDigital
TMHSDigital merged commit 33078b6 into main Sep 22, 2026
9 checks passed
@TMHSDigital
TMHSDigital deleted the docs/readme-for-a-stranger branch September 22, 2026 02:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant