docs(protocols): add P11, the benchmark protocol - #596
Open
breimanntools wants to merge 1 commit into
Open
breimanntools wants to merge 1 commit into
breimanntools wants to merge 1 commit into
Conversation
…cores The package ships fourteen benchmark tables across three prediction levels but no procedure over them. Two scores computed a month apart, on a different slice with a different split and seed, were therefore not comparable, so a performance regression was invisible and "method X improves over baseline" could not be checked. The data existed; the protocol over it did not. P11 pins the four decisions a comparable score needs, and records the result: - dataset: AA_CASPASE3 (residue), DOM_GSEC (domain), SEQ_AMYLO (protein), the trio the rest of the documentation already teaches on. The residue set is selected explicitly -- 20 seeded proteins carrying at least three sites, all their sites plus 25 seeded non-sites -- because load_dataset's deterministic head-of-class selection returns every negative from a single protein at small n, which is one protein, not a benchmark set. - split: StratifiedGroupKFold(5, shuffle=True, random_state=42) bound with aa.bind_groups on the protein accession, so windows cut from one protein cannot straddle a fold, scored by the pooled out-of-fold principle. - metric: ROC-AUC at all three levels against the aac and dpc composition baselines through identical folds (AAPred.eval's baseline= mode), plus per-protein average precision at the residue level. - seed: 42 throughout. random=False is pinned deliberately: load_dataset takes no random_state and aa.options["random_state"] does not reach it, so random=True is not reproducible and cannot appear in a benchmark. The tolerance is derived, not declared. A ten-seed sweep measures a widest across-seed standard deviation of 0.014, so three sigma rounds up to the 0.05 absolute ROC-AUC that separates noise from a regression. A tolerance picked a priori is either so tight it fires on unrelated changes or so loose it hides a real drop. Rerunning the pinned procedure reproduces every score exactly (max absolute difference 0.0), which is asserted in the notebook rather than claimed in prose. Two concept figures carry the teaching. The fold-composition panel shows the ungrouped split putting all 20 proteins in both halves of every fold while the grouped split holds 16 against 4, and reports honestly that the score gap on this set is small: you cannot know that until you measure it. The seed panel shows the free-seed cloud the tolerance is read off. The scoreboard reports that the amino-acid composition baseline beats the CPP feature set at the protein level. That is the protocol working: a benchmark that could only confirm the house method would be worth nothing. No new public symbols. The deliverable is the documented procedure, built from load_dataset, CPP, aa.bind_groups, AAPred.eval and comp_per_protein_ap as they already ship. Building the runner that automates it stays a separate concern. Co-Authored-By: Claude Fable 5.1 <[email protected]>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #596 +/- ##
=======================================
Coverage 95.39% 95.39%
=======================================
Files 222 222
Lines 23387 23387
Branches 4073 4073
=======================================
Hits 22309 22309
Misses 631 631
Partials 447 447 see 2 files with indirect coverage changes
🚀 New features to boost your workflow:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
protocols/protocol11_benchmark.ipynb— a reproducible evaluation protocol across the three prediction levels, assembled entirely from what already ships:load_dataset,SequenceFeature,CPP,aa.bind_groups,AAPred.eval(cv=, baseline=),comp_per_protein_ap.Zero new public symbols, as intended: #25 is the protocol, #488 is the runner that would later automate it, and the runner is not built here.
The four pinned decisions
AA_CASPASE3(residue),DOM_GSEC(domain),SEQ_AMYLO(protein): the trio the rest of the docs already teach on.StratifiedGroupKFold(5, shuffle=True, random_state=42)bound viaaa.bind_groupson the protein accession, scored out-of-fold. A benchmark that ignored grouping would teach the very leak this package just fixed.aac/dpccomposition baselines through identical folds, plus per-protein AP at the residue level.random=Falsepinned.Recorded baselines (seed 42): residue cpp 0.95 / aac 0.85 / dpc 0.87 · domain cpp 0.93 / 0.74 / 0.62 · protein cpp 0.83 / aac 0.91 / dpc 0.90.
The tolerance is derived, not declared: a ten-seed sweep measured a widest across-seed σ of 0.014, so 3σ rounds to 0.05 absolute ROC-AUC. Reproducibility is asserted in the notebook — rerun max absolute difference 0.0.
Three findings it reports honestly
load_dataset(random=True)is unseeded, so the protocol pinsrandom=False. Now fixed in fix(data): seed load_dataset's random sampling with random_state #593 / filed as load_dataset is not deterministic: set-ordered class blocks, and a regex character class that can gap canonical residues #588.random=Falsealone does not give a usable residue benchmark: entries readCASPASE3_<protein>_pos<index>and the deterministic head-of-class selection puts all 30 negatives in one protein atn=30. The protocol therefore defines its own seeded selection rule — and defining that rule is the work the issue asked for.df_folds_reveals the ungrouped split putting all 20 proteins in both halves of every fold versus 16-vs-4 grouped.Verification
pytest --nbmake protocols/→ 11 passed in 175s (the whole set stayed green). Docs gate:0 error · 167 critical, baseline held. 341api_tests. The captured Sphinx log shows no warning mentioning the new pages.Two things for you
notebook_examples.yml,CONTRIBUTING.rst+ its docs copy, and.claude/rules/workflow.md. Two are CONFIRM-FIRST, so none were touched; they want one approval together.Refs #25.
🤖 Generated with Claude Code