docs(readme): rewrite for someone who arrived from a link - #14
Merged
Merged
Conversation
The README was written by someone who already knew why the project exists, and it showed. 195 words passed before the tool was positively described, and four terms appeared as if common knowledge: Jev, JevBench, Benchmark Heaven and cascade. "Decision model", the class of thing this tool measures, was never defined at all. What changed. A "What it is" section now precedes "What this is not", which stays where it was. You cannot appreciate "not a leaderboard" before knowing what the thing is, and the not-a-leaderboard framing was previously the entire first screen. Decision model, calibration, cascade and Noul are each defined in a sentence where they first appear, and the Jev wire format gets a gloss in the adapters section. ECE is explained as a procedure before its floor is discussed, and binning noise is named rather than assumed. The argument now leads with the real report line as its evidence instead of paraphrasing the same figures as prose bullets and quoting them again sixty lines later. The figures appeared twice in two formats; the version with authority was the less prominent one. Quickstart states its prerequisites, says plainly that this is not on PyPI and that cloning is the install path, and offers bash alongside PowerShell. CI proves Ubuntu works and most people who would use a Python evaluation tool are not on Windows. It also documents `uv sync --extra local`, without which the local arm fails every case, and it no longer tells the reader to pass `--pricing my-pricing.json`, a file that never existed. It shows how to make one. Three overstatements corrected. "Adding a vendor is config, not code" is labeled as intent verified against one endpoint, linking #11. The adapters table gains a "Run for real" column, because two of the three transports have never left the test suite and presenting four rows as equally available was an omission that functioned as a claim. "Windows is the only verified platform" was stale: CI covers Ubuntu too. Adds a maturity note near the top: v0.1.0, one maintainer, the API will change in v0.2. The GitHub description is aligned with the opening line. The example report excerpt no longer leads with the cost line, which read oddly: a seeded mock obviously has no price, so it demonstrated nothing. The three quotes now all show the tool declining to answer, which is the point. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Everything from the Part 3 review, plus the six additions.
The review's central finding was that 195 words passed before the tool was
positively described, and four terms appeared as if common knowledge:
Jev,JevBench,Benchmark Heaven,cascade. "Decision model", the class of thingthis tool measures, was never defined at all.
Structural
You cannot appreciate "not a leaderboard" before knowing what the thing is,
and that framing was previously the entire first screen.
paraphrasing the same figures as prose bullets and quoting them again sixty
lines later. The figures appeared twice in two formats and the version with
authority was the less prominent one.
proves Ubuntu works, and most people who would use a Python evaluation tool
are not on Windows.
Definitions added
decision model,calibration,cascade,Noul,binning noise, and theJev wire format. ECE is now explained as a procedure before its floor isdiscussed, because the floor argument is unreadable if you do not know what is
being floored.
Overstatements corrected
against one endpoint, linking Add an adapter config for a second Jev-compatible endpoint #11.
transports have never left the test suite (Run the local arm for real: local_logits and generative have never left the test suite #3), and presenting four rows as
equally available was an omission that functioned as a claim.
From the independent review of the rendered page
--pricing my-pricing.jsonpointed at a file that never existed. Anyonecopying it failed on their first paid run. It now shows
cp docs/pricing.example.json my-pricing.json.python --version, since 3.11 gives a resolver error rather than a message.v0.2.
uv sync --extra localdocumented. Without it the local arm fails everycase.
read oddly: a seeded mock obviously has no price, so it demonstrated nothing.
All three quotes now show the tool declining to answer.
Kept on your instruction
"Trustworthy" in the opening line, as a fair summary of calibration plus
discrimination. The GitHub description is aligned to match it.
Verified
All four block quotes checked whitespace-normalised against the committed
report: verbatim. Both internal anchors resolve. No em dashes, no emoji. Gate
green.