Skip to content

Add discrimination and classwise calibration; fix the mislabelled line, the gold count and two edge cases - #8

Merged
rlaope merged 7 commits into
mainfrom
fix/report-bugs
Sep 26, 2026
Merged

rlaope merged 7 commits into
mainfrom
fix/report-bugs

Conversation

@rlaope

@rlaope rlaope commented Sep 26, 2026

Copy link
Copy Markdown
Owner

Summary

Discrimination: asks whether confidence ranks the model's own errors.

  • The statistics are in jeval/calibration.py: AUROC, the risk-coverage curve and AURC, with bootstrap intervals.
  • A new report section and a terminal line per question.
  • No interval is reported when a sample has fewer than 20 wrong or right answers; the reading says so instead.

Classwise calibration for choice questions:

  • Per-class ECE, with the class a cost action fires on marked.
  • Truncated probability maps are refused and counted.
  • New class_inflated synth mode. Existing modes still produce byte-identical output.

Fixes:

  • The reliability chart drew the recommended threshold as "line in use". It now draws both lines and names each.
  • The data-quality gold count included score answers (767 instead of the 696 actually measured).
  • sweep(n_boot=0) crashed.
  • The first tab stayed highlighted after choosing another.

Test plan

  • Synthetic tests were written first for both features.
  • Independent verification:
    • AUROC, risk-coverage, AURC and classwise ECE match brute force on 3,000 random cases.
    • Tolerances were checked (±0.045 ≈ 4 SE).
    • README figures reproduce.
  • 19 mutations were run. The blocker found and the 5 behaviours no test caught now have tests:
    • null with few wrong answers
    • interval widening
    • absent class read as 0
    • gold-only discrimination
    • the classwise "clears" branch
    • noul discrimination on the stated-probability scale
    • the tab fix
  • uv run pytest, ruff and mypy are clean; goldens, examples and README screenshots regenerated.

Add AUROC (Mann-Whitney, ties at half weight), the risk-coverage curve
and its area (AURC), each with a percentile bootstrap interval, and a
one-line reading that says when confidence ranks nothing.

Add classwise ECE for choice questions: every class is measured on its
own probability over records whose map sums to 1 within 0.02; truncated
or missing maps are refused and counted, and a class with fewer than 30
labels is refused and says so.

Add a class_inflated synthetic mode: confidence is inflated only when
one class is predicted and the other classes compensate, so top-1 ECE
stays near zero while that class's classwise ECE does not. Existing
modes keep byte-identical output.

Signed-off-by: rlaope <[email protected]>
A new Discrimination section follows Reliability: one risk-coverage
chart per question (coverage = share of decisions automated, risk =
error rate among them) against the no-ranking line and the perfect
ranking, with the recommended line marked in the accent where a cost
action sets one, AUROC and AURC with intervals as key figures, and a
plain reading that says when confidence ranks nothing.

A choice question's reliability block gains a per-class table: class,
labels, ECE, interval, and which cost action fires on the class, with
the worst class named in one sentence and truncated maps counted.

jeval report prints AUROC per question.

Signed-off-by: rlaope <[email protected]>
The first tab is rendered with is-active for a reader without
JavaScript, and the tab script only updated aria-selected, so after a
click the first tab stayed highlighted beside the chosen one.

Signed-off-by: rlaope <[email protected]>
Regenerate the example report and the workbench from it, quote the
intent question's AUROC reading and per-class ECEs as the committed
artifact states them, update the file size to 340 KB, and guard the
quoted figures in tests/test_readme_example.py.

Signed-off-by: rlaope <[email protected]>
Signed-off-by: rlaope <[email protected]>

# Conflicts:
#	CHANGELOG.md
#	jeval/calibration.py
…gold as measured

The reliability chart drew the recommended threshold as "line in use"; it now draws both lines,
each named, and the block states both. The data-quality gold row counted score answers and read 767
beside a report measured on 696; it counts the measured sample. sweep(n_boot=0) crashed; it reports
no interval. Each is pinned by a test; goldens, examples and README screenshots are regenerated.

Signed-off-by: rlaope <[email protected]>
Independent verification found the percentile interval excluding 0.5 up to 42% of the time under a
null with two or three wrong answers, and the reading calling it separation. Below 20 wrong (or right)
answers — the report's bin floor — there is no interval and the reading says the sample is too thin;
a null synthetic test pins it. Tests also pin five behaviours mutation testing showed were unguarded:
the interval widened to contain its estimate, an absent class read as zero, gold-only
discrimination, the classwise 'clears' branch, and a yes/no question ranked on its stated
probability, plus the tab highlight fix. The perfect-ranking line is dotted, and coverage levels print
three digits.

Signed-off-by: rlaope <[email protected]>
@rlaope
rlaope merged commit 7568620 into main Sep 26, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant