Skip to content

feat(prediction): audit an evaluation setup for leakage risks - #592

Open
breimanntools wants to merge 1 commit into
masterfrom
feat/479-audit-leakage
Open

breimanntools wants to merge 1 commit into
masterfrom
feat/479-audit-leakage

Conversation

@breimanntools

Copy link
Copy Markdown
Owner

Most evaluation leaks are invisible until someone notices an implausibly high score. aa.audit_leakage inspects a dataset and a split and returns a plain table of findings.

df_audit = aa.audit_leakage(df_seq=df_seq, groups=df_seq["entry"],
                            splits=cv.split(X, labels), raise_on="high")
df_audit.attrs["status"]   # 'ok' | 'low' | 'medium' | 'high'

Eight checks: duplicate sequences, train/test overlap, duplicates across folds, the same protein in several folds, group overlap, a target-derived feature, fold-size anomaly, class-balance anomaly. Dataset-level checks run without splits; fold-level checks compare train against test within each fold.

Returns a plain 4-column DataFrame (check, severity, detail, ids), worst first, with the verdict in df.attrs["status"] — no bespoke result type, which the issue explicitly forbids. Pairs with bind_groups: splits=bind_groups(...).split(X, y) is the natural input.

KPIs

Asserted both ways: a duplicate placed in two folds produces a high finding naming the ids; a clean setup returns an empty table with status == "ok"; raise_on="high" raises only when a high finding exists.

Verification

64 tests (44 + 20 across the two house classes), hypothesis on 4, error-message match= throughout. 1189 tests pass across prediction_tests + api_tests. pyright 0 errors. Docstring checker 0 defects, drift 0. Docs gate at baseline (167). Notebook: 7/7 public params by name, no markdown headings.

Two things to look at

  1. Findings are aggregated per check, with the affected folds named in detail, rather than one row per fold — a 10-fold leak would otherwise emit 10 near-identical rows. That trades machine-addressable fold ids for readability; it is the one genuine design fork here.
  2. Thresholds are fixed, not parameters (|r| >= 0.95, fold-size ratio 2.0, class-share deviation 0.2, ids capped at 10), so a report means the same thing everywhere. Documented in Notes.

One test caught a true positive that is worth knowing: GroupKFold over proteins whose label is the protein's parity yields single-class folds, and the audit correctly reports it. The test expectation was fixed, not the code, and the notebook explains it.

Refs #479 (no closing keyword — shipped code is a first draft).

🤖 Generated with Claude Code

Leakage inflates a score without ever raising an error, and on small windowed
datasets it is easy to introduce and hard to see. aa.audit_leakage runs a set of
cheap heuristics over whatever parts of a setup are supplied and returns a plain
DataFrame of findings (check, severity, detail, ids), worst first, with the
overall verdict in df_audit.attrs["status"] so no bespoke result type is needed.

Without splits it audits the dataset, which is the check to run before choosing a
split; with splits it also compares the training against the test part within each
fold. It reports duplicate sequences, a row index, protein or group present on
both sides of a fold, a feature correlating with the label at |r| >= 0.95, an
empty or strongly uneven fold, and a single-class or skewed fold.

Windowed input is handled correctly: the sampler emits both the repeated parent
'sequence' and the 'window' cut from it, so the window takes precedence and
sibling windows of one protein are not mistaken for duplicate samples.

raise_on turns findings at or above a severity into a bare ValueError naming
them; the default reports only and never raises, so an audit can be dropped into
a workflow without changing its control flow. The checks are heuristic and a
clean report is not a proof that no leakage exists; the severities are labels for
a human reader, not a machine taxonomy, since whether a finding must block a
workflow is a policy decision left to the caller.

Pairs with bind_groups, which prevents the group leak this reports: a split from
bind_groups(...).split(X, y) feeds straight into splits=.

Refs #479

Co-Authored-By: Claude Fable 5.1 <[email protected]>
@codecov

codecov Bot commented Sep 18, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.87926% with 23 lines in your changes missing coverage. Please review.
✅ Project coverage is 95.35%. Comparing base (3c11c90) to head (dc1ea56).
⚠️ Report is 10 commits behind head on master.

Files with missing lines Patch % Lines
aaanalysis/prediction/_backend/audit_leakage.py 92.85% 7 Missing and 6 partials ⚠️
aaanalysis/prediction/_audit_leakage.py 91.73% 6 Missing and 4 partials ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##           master     #592      +/-   ##
==========================================
- Coverage   95.39%   95.35%   -0.04%     
==========================================
  Files         222      224       +2     
  Lines       23387    23710     +323     
  Branches     4073     4146      +73     
==========================================
+ Hits        22309    22609     +300     
- Misses        631      644      +13     
- Partials      447      457      +10     
Files with missing lines Coverage Δ
aaanalysis/__init__.py 92.85% <ø> (ø)
aaanalysis/_constants.py 100.00% <100.00%> (ø)
aaanalysis/prediction/__init__.py 100.00% <100.00%> (ø)
aaanalysis/prediction/_audit_leakage.py 91.73% <91.73%> (ø)
aaanalysis/prediction/_backend/audit_leakage.py 92.85% <92.85%> (ø)

... and 1 file with indirect coverage changes

Components Coverage Δ
cpp_core 95.97% <ø> (ø)
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant