Skip to content

feat(prediction): bind group labels to a splitter for leak-free evaluation - #580

Merged
breimanntools merged 2 commits into
masterfrom
feat/478-bind-groups
Sep 18, 2026
Merged

breimanntools merged 2 commits into
masterfrom
feat/478-bind-groups

Conversation

@breimanntools

Copy link
Copy Markdown
Owner

Protein datasets are full of dependent samples: several windows cut from one protein, near-identical homologues across a family. A plain split scatters them over train and test, so the model is scored partly on what it already saw.

A scikit-learn group splitter prevents that — but it needs a groups array at split time, and every cross_val_* call in this package is made without one, so a stock GroupKFold cannot be used at all. bind_groups closes that gap by binding the labels to the splitter, which is why AAPred.eval(cv=...) and ModelEvaluator.run(cv=...) accept a group-aware splitter with no change to either — the issue's headline acceptance criterion.

Why one adapter rather than three named classes

It wraps any splitter, so GroupKFold, StratifiedGroupKFold (keeps the class balance too), LeaveOneGroupOut (leave-one-protein-out, or leave-one-cluster-out over homology clusters) and GroupShuffleSplit are all covered by sklearn's own tested implementations. The domain vocabulary lives in what is passed as groups, which is also what makes the same call work for proteins, families and externally computed clusters. Two of the three names the issue proposed (LeaveOneProteinOut, LeaveOneClusterOut) would have been the same implementation.

cv = aa.bind_groups(StratifiedGroupKFold(n_splits=5), groups=df_seq["entry"])
df_eval = aap.eval(X, labels=labels, cv=cv)     # eval unchanged
cv.df_folds_                                     # group sizes + class balance per fold

Two guards

  • Every fold is verified. A splitter that ignores groups (e.g. KFold) raises a ValueError naming the shared groups, instead of leaking silently. allow_overlap=True permits it deliberately, e.g. to reproduce an ungrouped baseline.
  • An infeasible fold count is rejected at construction, with the requested folds and the group count in the message.

Once the folds are consumed, df_folds_ reports per-fold sample counts, group counts and class balance, so the uneven folds a group splitter necessarily produces are inspectable rather than assumed.

The leak, measured

On a seeded fixture of twelve proteins with six near-identical windows each, where the label is a property of the protein:

split accuracy
random (leaky) 1.00
protein-grouped 0.67

That contrast is the example notebook.

Verification

  • 43 unit tests in tests/unit/prediction_tests/test_bind_groups.py, including the no-overlap KPI over all folds, the infeasible-fold message, a hand-computed pos_rate_test golden value, reproducibility of a seeded splitter, and the AAPred.eval end-to-end path.
  • 1237 tests pass across prediction_tests + api_tests + config_tests.
  • Docstring checker and doc-vs-signature drift: 0 defects.
  • The new docs gate passes at baseline (0 errors, 167 critical). It caught this branch first: the example notebook originally used ## headings, which become CRITICAL: Unexpected section title in the generated page that is include::d into the docstring. Existing example notebooks carry no headings; this one now matches.
  • [Roberts17] added to the references (cross-validation for structured data).

Addresses #478.

🤖 Generated with Claude Code

breimanntools and others added 2 commits September 18, 2026 15:47
…ation

Protein datasets are full of dependent samples: several windows cut from
one protein, near-identical homologues across a family. A plain split
scatters them over train and test, so the model is scored partly on what
it already saw.

A scikit-learn group splitter prevents that, but it needs a `groups`
array at split time, and every cross_val_* call in this package is made
without one -- so a stock GroupKFold cannot be used at all. bind_groups
closes that gap by binding the labels to the splitter, which is why
AAPred.eval(cv=...) and ModelEvaluator.run(cv=...) take a group-aware
splitter with no change to either.

One adapter rather than three named classes: it wraps any splitter, so
GroupKFold, StratifiedGroupKFold (keeps the class balance too),
LeaveOneGroupOut (leave-one-protein-out, or leave-one-cluster-out over
homology clusters) and GroupShuffleSplit are all covered by sklearn's own
tested implementations, and the vocabulary lives in what is passed as
`groups`.

Two guards the issue asked for:

- every fold is verified, so a splitter that ignores groups (KFold) is a
  ValueError naming the shared groups rather than a silent leak;
  allow_overlap=True permits it deliberately
- an infeasible fold count is rejected at construction, with the
  requested folds and the group count in the message

Once the folds are consumed, df_folds_ reports per-fold sample counts,
group counts and class balance, so uneven group folds are inspectable
rather than assumed.

Measured on a seeded fixture of twelve proteins with six near-identical
windows each: a random split reports 1.00 accuracy, the protein-grouped
split 0.67. That contrast is the example notebook.

The package performs the split and reports its properties; it runs no
homology search or clustering itself.

Addresses #478.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
The pyright ratchet caught 14 new diagnostics: every helper declared its
required parameters as `=None`, so `None` propagated into every inference
downstream (`self._cv.split` on `None`, `len(self._groups)` on `None`).

Required parameters now carry no default, which is the rule the earlier
burn-down established. No behaviour change: the call sites already passed
every argument by keyword, and `_comp_pos_rate` keeps `labels=None` because
None is a real value there (no labels -> NaN rate).

Co-Authored-By: Claude Fable 5.1 <[email protected]>
@codecov

codecov Bot commented Sep 18, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.85714% with 6 lines in your changes missing coverage. Please review.
✅ Project coverage is 95.39%. Comparing base (b4d81d6) to head (b13538d).

Files with missing lines Patch % Lines
aaanalysis/prediction/_bind_groups.py 92.10% 4 Missing and 2 partials ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##           master     #580      +/-   ##
==========================================
- Coverage   95.39%   95.39%   -0.01%     
==========================================
  Files         221      222       +1     
  Lines       23303    23387      +84     
  Branches     4057     4073      +16     
==========================================
+ Hits        22231    22309      +78     
- Misses        627      631       +4     
- Partials      445      447       +2     
Files with missing lines Coverage Δ
aaanalysis/__init__.py 92.85% <ø> (ø)
aaanalysis/_constants.py 100.00% <100.00%> (ø)
aaanalysis/prediction/__init__.py 100.00% <100.00%> (ø)
aaanalysis/prediction/_bind_groups.py 92.10% <92.10%> (ø)
Components Coverage Δ
cpp_core 95.97% <ø> (ø)
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@breimanntools
breimanntools merged commit 3c11c90 into master Sep 18, 2026
27 checks passed
@breimanntools
breimanntools deleted the feat/478-bind-groups branch September 18, 2026 17:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant