Skip to content

fix(report): choose the cascade threshold on one half and score it on the other - #81

Merged
TMHSDigital merged 1 commit into
mainfrom
fix/cascade-on-held-out-rows
Sep 23, 2026
Merged

TMHSDigital merged 1 commit into
mainfrom
fix/cascade-on-held-out-rows

Conversation

@TMHSDigital

Copy link
Copy Markdown
Owner

Fixes #30.

The problem. The cascade chose its threshold and reported its coverage and cost on the same rows, so the printed figures were its in-sample best case. With a temperature recommended those rows were the eval split. With none, it was every row, while the text still said "held-out". On 300 synthetic rows with recalibration refused, it printed threshold 0.943 ... over 300 rows, covered accuracy 1.000.

What changed

  • The threshold is now chosen on the fit half and scored on the held-out half. That is the recalibration split when there is one, or a split made the same way (make_split, same fraction and seed) when there is not.
  • When a temperature was recommended, both halves are on the recalibrated scale, so the held-out rows are unseen by the temperature and by the threshold.
  • The 200-row minimum counts held-out rows, so a threshold now needs at least 400 scored rows. This changes which runs print a threshold, and the cost and coverage printed with it.
  • The report says where the threshold was chosen and where it was scored ("chosen on 300 rows and scored on 300 held-out rows"). PLAN says the same.

Checked

  • New tests failed on main and pass now:
    • 300 calibrated rows (no temperature, 150 held out) now refuse a threshold and say why;
    • 600 rows name where the threshold was chosen and where it was scored.
  • Every existing cascade test still passes, including the escalate-everything wording from fix(cascade): consider escalating everything, and price correct answers over priced rows #80.
  • ruff, mypy --strict, the full pytest suite, and the site build all pass. The example report supplies no costs, so it is unchanged.

🤖 Generated with Claude Code

… the other

The cascade chose its threshold and reported its coverage and cost on the
same rows, so the figures were its in-sample best case. With a temperature
recommended those were the eval rows; with none, every row, while the text
still said "held-out". The 200-row minimum counted the same rows.

The threshold is now chosen on the fit half and scored on the held-out
half, using the recalibration split when there is one and a split made the
same way when there is not. Both halves are on the recalibrated scale when
a temperature was recommended, so the held-out rows are unseen by the
temperature and the threshold alike. The 200-row minimum counts held-out
rows. The report says where the threshold was chosen and where it was
scored. PLAN says the same.

Fixes #30.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@TMHSDigital
TMHSDigital merged commit 64f7f8e into main Sep 23, 2026
17 checks passed
@TMHSDigital
TMHSDigital deleted the fix/cascade-on-held-out-rows branch September 23, 2026 18:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The cascade picks its threshold on the same rows it reports on

1 participant