Add discrimination and classwise calibration; fix the mislabelled line, the gold count and two edge cases - #8
Merged
Merged
Conversation
Add AUROC (Mann-Whitney, ties at half weight), the risk-coverage curve and its area (AURC), each with a percentile bootstrap interval, and a one-line reading that says when confidence ranks nothing. Add classwise ECE for choice questions: every class is measured on its own probability over records whose map sums to 1 within 0.02; truncated or missing maps are refused and counted, and a class with fewer than 30 labels is refused and says so. Add a class_inflated synthetic mode: confidence is inflated only when one class is predicted and the other classes compensate, so top-1 ECE stays near zero while that class's classwise ECE does not. Existing modes keep byte-identical output. Signed-off-by: rlaope <[email protected]>
A new Discrimination section follows Reliability: one risk-coverage chart per question (coverage = share of decisions automated, risk = error rate among them) against the no-ranking line and the perfect ranking, with the recommended line marked in the accent where a cost action sets one, AUROC and AURC with intervals as key figures, and a plain reading that says when confidence ranks nothing. A choice question's reliability block gains a per-class table: class, labels, ECE, interval, and which cost action fires on the class, with the worst class named in one sentence and truncated maps counted. jeval report prints AUROC per question. Signed-off-by: rlaope <[email protected]>
The first tab is rendered with is-active for a reader without JavaScript, and the tab script only updated aria-selected, so after a click the first tab stayed highlighted beside the chosen one. Signed-off-by: rlaope <[email protected]>
Regenerate the example report and the workbench from it, quote the intent question's AUROC reading and per-class ECEs as the committed artifact states them, update the file size to 340 KB, and guard the quoted figures in tests/test_readme_example.py. Signed-off-by: rlaope <[email protected]>
Signed-off-by: rlaope <[email protected]> # Conflicts: # CHANGELOG.md # jeval/calibration.py
…gold as measured The reliability chart drew the recommended threshold as "line in use"; it now draws both lines, each named, and the block states both. The data-quality gold row counted score answers and read 767 beside a report measured on 696; it counts the measured sample. sweep(n_boot=0) crashed; it reports no interval. Each is pinned by a test; goldens, examples and README screenshots are regenerated. Signed-off-by: rlaope <[email protected]>
Independent verification found the percentile interval excluding 0.5 up to 42% of the time under a null with two or three wrong answers, and the reading calling it separation. Below 20 wrong (or right) answers — the report's bin floor — there is no interval and the reading says the sample is too thin; a null synthetic test pins it. Tests also pin five behaviours mutation testing showed were unguarded: the interval widened to contain its estimate, an absent class read as zero, gold-only discrimination, the classwise 'clears' branch, and a yes/no question ranked on its stated probability, plus the tab highlight fix. The perfect-ranking line is dotted, and coverage levels print three digits. Signed-off-by: rlaope <[email protected]>
4 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Discrimination: asks whether confidence ranks the model's own errors.
jeval/calibration.py: AUROC, the risk-coverage curve and AURC, with bootstrap intervals.Classwise calibration for
choicequestions:class_inflatedsynth mode. Existing modes still produce byte-identical output.Fixes:
sweep(n_boot=0)crashed.Test plan
uv run pytest,ruffandmypyare clean; goldens, examples and README screenshots regenerated.