Skip to content

Say each finding as a sentence and draw it; fold the tables underneath - #9

Merged
rlaope merged 9 commits into
mainfrom
design/figures-first
Sep 26, 2026
Merged

rlaope merged 9 commits into
mainfrom
design/figures-first

Conversation

@rlaope

@rlaope rlaope commented Sep 26, 2026

Copy link
Copy Markdown
Owner

Builds on #8. Merge #8 first; after that, this diff shows only its own three commits.

Summary

Every report section now opens with its finding as a sentence and draws it. The table it came from is folded under "The numbers behind this, as a table", and every number the README quotes is still in it.

  • Cost: "Moving the line from 0.60 to 0.75 saves KRW 3,770,492 a month". Below it, four figure cards, each showing its new figure with the old one struck through.
  • Reliability: the ranges where what the model said and how often it was right part ways, drawn as said-versus-right rows.
  • Segment splits: one verdict card per action.
  • Label planning: progress cards.
  • Data quality: one composition bar.
  • Score bands and per-class calibration: rows and chips.
  • Type size: larger throughout — 17px body, 40px headline figures, 12.5–13px chart labels.

Test plan

  • uv run pytest, ruff and mypy are clean. Goldens were regenerated for the larger chart labels.
  • Checked visually with Playwright at 1280px in light and dark, and at 390px mobile, where the figure cards sit two to a row.
  • Every new SVG has <title> and <desc>. The composition bar and the rows carry their meaning in words and patterns as well as colour.
  • The README size claim (374 KB), the screenshots and the guard tests were updated.

Add AUROC (Mann-Whitney, ties at half weight), the risk-coverage curve
and its area (AURC), each with a percentile bootstrap interval, and a
one-line reading that says when confidence ranks nothing.

Add classwise ECE for choice questions: every class is measured on its
own probability over records whose map sums to 1 within 0.02; truncated
or missing maps are refused and counted, and a class with fewer than 30
labels is refused and says so.

Add a class_inflated synthetic mode: confidence is inflated only when
one class is predicted and the other classes compensate, so top-1 ECE
stays near zero while that class's classwise ECE does not. Existing
modes keep byte-identical output.

Signed-off-by: rlaope <[email protected]>
A new Discrimination section follows Reliability: one risk-coverage
chart per question (coverage = share of decisions automated, risk =
error rate among them) against the no-ranking line and the perfect
ranking, with the recommended line marked in the accent where a cost
action sets one, AUROC and AURC with intervals as key figures, and a
plain reading that says when confidence ranks nothing.

A choice question's reliability block gains a per-class table: class,
labels, ECE, interval, and which cost action fires on the class, with
the worst class named in one sentence and truncated maps counted.

jeval report prints AUROC per question.

Signed-off-by: rlaope <[email protected]>
The first tab is rendered with is-active for a reader without
JavaScript, and the tab script only updated aria-selected, so after a
click the first tab stayed highlighted beside the chosen one.

Signed-off-by: rlaope <[email protected]>
Regenerate the example report and the workbench from it, quote the
intent question's AUROC reading and per-class ECEs as the committed
artifact states them, update the file size to 340 KB, and guard the
quoted figures in tests/test_readme_example.py.

Signed-off-by: rlaope <[email protected]>
Signed-off-by: rlaope <[email protected]>

# Conflicts:
#	CHANGELOG.md
#	jeval/calibration.py
…gold as measured

The reliability chart drew the recommended threshold as "line in use"; it now draws both lines,
each named, and the block states both. The data-quality gold row counted score answers and read 767
beside a report measured on 696; it counts the measured sample. sweep(n_boot=0) crashed; it reports
no interval. Each is pinned by a test; goldens, examples and README screenshots are regenerated.

Signed-off-by: rlaope <[email protected]>
Independent verification found the percentile interval excluding 0.5 up to 42% of the time under a
null with two or three wrong answers, and the reading calling it separation. Below 20 wrong (or right)
answers — the report's bin floor — there is no interval and the reading says the sample is too thin;
a null synthetic test pins it. Tests also pin five behaviours mutation testing showed were unguarded:
the interval widened to contain its estimate, an absent class read as zero, gold-only
discrimination, the classwise 'clears' branch, and a yes/no question ranked on its stated
probability, plus the tab highlight fix. The perfect-ranking line is dotted, and coverage levels print
three digits.

Signed-off-by: rlaope <[email protected]>
…ded underneath

Larger type (17px body, 40px headline figures). The cost section opens with what moving the line buys
and four figure cards; reliability names the ranges that miss as said-versus-right rows; segment
splits become one verdict card per action; label planning becomes progress cards; data quality
becomes one composition bar; score bands and per-class calibration become rows and chips. Every
table stays, folded under 'The numbers behind this, as a table'.

Signed-off-by: rlaope <[email protected]>
The example report, workbench and README screenshots are rebuilt; the impact screenshot now shows
the finding and its figure cards, the example is 374 KB, and the design language in examples/README
describes findings before tables. A README guard matches the classwise block by its class token.

Signed-off-by: rlaope <[email protected]>
@rlaope
rlaope merged commit c077b4d into main Sep 26, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant