Skip to content

fix(cascade): consider escalating everything, and price correct answers over priced rows - #80

Merged
TMHSDigital merged 1 commit into
mainfrom
fix/cascade-escalate-all-and-cost-per-correct
Sep 23, 2026
Merged

TMHSDigital merged 1 commit into
mainfrom
fix/cascade-escalate-all-and-cost-per-correct

Conversation

@TMHSDigital

Copy link
Copy Markdown
Owner

Fixes #34. Fixes #35.

#34: the cascade never considered escalating every case.

  • The candidate thresholds were 0 plus every observed score, and a row is covered when its score reaches the threshold. So the highest score still covered the rows that reached it, and "escalate everything" was never a candidate.
  • With every case wrong, an error at $100 and an escalation at $1, it chose threshold 0.9 at $201 where escalating all three cost $3.
  • ESCALATE_ALL (infinity) is now the last candidate. When it wins, the report says in words that escalating every case is cheapest, instead of printing a threshold of inf.

#35: per_correct_usd understated the cost of a correct answer.

  • It divided the total over priced rows by correct answers across all rows: summarize([1.0, None], [True, True]) gave 0.5 where 1.0 is right.
  • It now counts correct answers among priced rows only.

Checked

  • New tests failed on main and pass now:
    • the all-wrong case now escalates all three at $3;
    • per_correct_usd is 1.0 in the case above;
    • the report says "Escalating every case is cheapest" and prints no inf. This test forces the optimum to the new candidate, because the mock refuses an accuracy below chance.
  • The test that pinned the exact list of candidate thresholds now includes ESCALATE_ALL.
  • ruff, mypy --strict, the full pytest suite, and the site build all pass.

🤖 Generated with Claude Code

…rs over priced rows

The cascade's candidate thresholds were 0 plus every observed score, and a
row is covered when its score reaches the threshold, so the highest score
still covered the rows that reached it: escalating every case was never a
candidate. With every case wrong, an error at $100 and an escalation at $1,
it chose 0.9 at $201 where escalating all three cost $3. ESCALATE_ALL
(infinity) is now the last candidate, and the report says in words that
escalating every case is cheapest rather than printing a threshold of inf.

CostSummary.per_correct_usd divided the total over priced rows by the
correct answers across all rows, so with unpriced rows it understated the
cost of a correct answer: summarize([1.0, None], [True, True]) gave 0.5
where 1.0 is right. It now counts correct answers among priced rows.

Fixes #34. Fixes #35.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@TMHSDigital
TMHSDigital merged commit df49c7f into main Sep 23, 2026
17 checks passed
@TMHSDigital
TMHSDigital deleted the fix/cascade-escalate-all-and-cost-per-correct branch September 23, 2026 18:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

per_correct_usd divides the priced-row cost by the correct rows across all rows The cascade never considers escalating every case

1 participant