Skip to content

docs(protocols): add P11, the benchmark protocol - #596

Open
breimanntools wants to merge 1 commit into
masterfrom
doc/25-benchmark-protocol
Open

breimanntools wants to merge 1 commit into
masterfrom
doc/25-benchmark-protocol

Conversation

@breimanntools

Copy link
Copy Markdown
Owner

protocols/protocol11_benchmark.ipynb — a reproducible evaluation protocol across the three prediction levels, assembled entirely from what already ships: load_dataset, SequenceFeature, CPP, aa.bind_groups, AAPred.eval(cv=, baseline=), comp_per_protein_ap.

Zero new public symbols, as intended: #25 is the protocol, #488 is the runner that would later automate it, and the runner is not built here.

The four pinned decisions

  • Datasets — AA_CASPASE3 (residue), DOM_GSEC (domain), SEQ_AMYLO (protein): the trio the rest of the docs already teach on.
  • Split — StratifiedGroupKFold(5, shuffle=True, random_state=42) bound via aa.bind_groups on the protein accession, scored out-of-fold. A benchmark that ignored grouping would teach the very leak this package just fixed.
  • Metric — ROC-AUC at all three levels against the aac / dpc composition baselines through identical folds, plus per-protein AP at the residue level.
  • Seed — 42 throughout, random=False pinned.

Recorded baselines (seed 42): residue cpp 0.95 / aac 0.85 / dpc 0.87 · domain cpp 0.93 / 0.74 / 0.62 · protein cpp 0.83 / aac 0.91 / dpc 0.90.

The tolerance is derived, not declared: a ten-seed sweep measured a widest across-seed σ of 0.014, so 3σ rounds to 0.05 absolute ROC-AUC. Reproducibility is asserted in the notebook — rerun max absolute difference 0.0.

Three findings it reports honestly

  1. load_dataset(random=True) is unseeded, so the protocol pins random=False. Now fixed in fix(data): seed load_dataset's random sampling with random_state #593 / filed as load_dataset is not deterministic: set-ordered class blocks, and a regex character class that can gap canonical residues #588.
  2. random=False alone does not give a usable residue benchmark: entries read CASPASE3_<protein>_pos<index> and the deterministic head-of-class selection puts all 30 negatives in one protein at n=30. The protocol therefore defines its own seeded selection rule — and defining that rule is the work the issue asked for.
  3. The grouping gap on this data is small (MCC 0.618 ungrouped → 0.576 grouped) and it says so rather than overselling. The concept figure shows the leak structurally instead: df_folds_ reveals the ungrouped split putting all 20 proteins in both halves of every fold versus 16-vs-4 grouped.

Verification

pytest --nbmake protocols/ → 11 passed in 175s (the whole set stayed green). Docs gate: 0 error · 167 critical, baseline held. 341 api_tests. The captured Sphinx log shows no warning mentioning the new pages.

Two things for you

  • The scores are post-selection optimistic: CPP features are selected once on the full set before the folds are cut. Kept deliberately — it freezes the feature set so a rerun compares methods rather than reshuffled features — and stated plainly in When to use it and Common mistakes. Running nested selection instead is a real design change and a much slower notebook.
  • "All ten protocols" is now stale in notebook_examples.yml, CONTRIBUTING.rst + its docs copy, and .claude/rules/workflow.md. Two are CONFIRM-FIRST, so none were touched; they want one approval together.

Refs #25.

🤖 Generated with Claude Code

…cores

The package ships fourteen benchmark tables across three prediction levels
but no procedure over them. Two scores computed a month apart, on a different
slice with a different split and seed, were therefore not comparable, so a
performance regression was invisible and "method X improves over baseline"
could not be checked. The data existed; the protocol over it did not.

P11 pins the four decisions a comparable score needs, and records the result:

- dataset: AA_CASPASE3 (residue), DOM_GSEC (domain), SEQ_AMYLO (protein), the
  trio the rest of the documentation already teaches on. The residue set is
  selected explicitly -- 20 seeded proteins carrying at least three sites, all
  their sites plus 25 seeded non-sites -- because load_dataset's deterministic
  head-of-class selection returns every negative from a single protein at
  small n, which is one protein, not a benchmark set.
- split: StratifiedGroupKFold(5, shuffle=True, random_state=42) bound with
  aa.bind_groups on the protein accession, so windows cut from one protein
  cannot straddle a fold, scored by the pooled out-of-fold principle.
- metric: ROC-AUC at all three levels against the aac and dpc composition
  baselines through identical folds (AAPred.eval's baseline= mode), plus
  per-protein average precision at the residue level.
- seed: 42 throughout. random=False is pinned deliberately: load_dataset takes
  no random_state and aa.options["random_state"] does not reach it, so
  random=True is not reproducible and cannot appear in a benchmark.

The tolerance is derived, not declared. A ten-seed sweep measures a widest
across-seed standard deviation of 0.014, so three sigma rounds up to the 0.05
absolute ROC-AUC that separates noise from a regression. A tolerance picked a
priori is either so tight it fires on unrelated changes or so loose it hides a
real drop. Rerunning the pinned procedure reproduces every score exactly (max
absolute difference 0.0), which is asserted in the notebook rather than
claimed in prose.

Two concept figures carry the teaching. The fold-composition panel shows the
ungrouped split putting all 20 proteins in both halves of every fold while the
grouped split holds 16 against 4, and reports honestly that the score gap on
this set is small: you cannot know that until you measure it. The seed panel
shows the free-seed cloud the tolerance is read off.

The scoreboard reports that the amino-acid composition baseline beats the CPP
feature set at the protein level. That is the protocol working: a benchmark
that could only confirm the house method would be worth nothing.

No new public symbols. The deliverable is the documented procedure, built from
load_dataset, CPP, aa.bind_groups, AAPred.eval and comp_per_protein_ap as they
already ship. Building the runner that automates it stays a separate concern.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
@codecov

codecov Bot commented Sep 18, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 95.39%. Comparing base (3c11c90) to head (4e53566).
⚠️ Report is 10 commits behind head on master.

Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff           @@
##           master     #596   +/-   ##
=======================================
  Coverage   95.39%   95.39%           
=======================================
  Files         222      222           
  Lines       23387    23387           
  Branches     4073     4073           
=======================================
  Hits        22309    22309           
  Misses        631      631           
  Partials      447      447           

see 2 files with indirect coverage changes

Components Coverage Δ
cpp_core 95.97% <ø> (ø)
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant