Skip to content

docs: compact the docstrings and fix the version stamps - #589

Open
breimanntools wants to merge 9 commits into
masterfrom
doc/docstrings-integration
Open

breimanntools wants to merge 9 commits into
masterfrom
doc/docstrings-integration

Conversation

@breimanntools

Copy link
Copy Markdown
Owner

The docstrings had absorbed the rationale of every feature as it was added, and the API pages opened with a stack of "Changed in version …" blocks before the reader learned what a method does. This is the cleanup, across four subpackages plus the version stamps.

720 insertions, 1440 deletions — net −720 lines, all of it docstring text.

What changed

before after
versionchanged 40 3
versionadded naming a pre-release (0.1.x) 68 0
CPP.run 215 lines 139
CPP.__init__ 118 79
ReliabilityModel 117 70
StructurePreprocessor.__init__ 128 41
SeqOpt (class) 70 31

The 3 surviving versionchanged are all 1.2.0, one line each, sitting under the parameter they describe — scikit-learn's placement. They are kept because a user hitting the next release needs them: ci changed from a percent to a fraction, so someone passing 90.0 would otherwise silently get wrong behaviour.

Version stamps keep the scikit-learn convention — every public class and function still carries exactly one .. versionadded::, and the checker rule is untouched. Only the floor moved: 68 directives naming 0.1.0 / 0.1.2 / 0.1.3 now read 1.0.0, since the public history starts at the first release and "added in 0.1.0" names a version no user installed. That is a deliberate convention, and the style guide and .claude/rules/docstrings.md both now say so.

What was cut, and what was kept

Cut: a 75-row normalization table in StructurePreprocessor.__init__ that duplicated feature_registry.NORMALIZATION_RECIPES; a 13-row SeqOpt→DEAP mapping (DEAP is a dev/test-only parity oracle the runtime never imports); CPP.run's 13-item df_feat column list, now a pointer to the generated contract page that a test keeps in sync; the O(batch_size × part_length × n_scales) memory prose; the "Why these defaults?" bootstrap essays; the _kws preamble.

Kept: the biology, and every load-bearing trap reduced to one plain sentence — splits auto-capping, the raw-PLM normalization trap in run_num, the sample-level mean_dif trap, the CPP-inside-each-fold leakage warning. Plus every **Experimental.** marker (17 before, 17 after), every [Key]_ citation, every .. include:: examples/*.rst, every named Returns.

Three docstrings that were factually wrong

  • AAPred claimed it "intentionally does not perform hyperparameter optimization" while fit(optimize_hyperparams=True) runs GridSearchCV
  • CPPPlot.feature_map's seq_char_fill carried "Now defaults to True"; the signature is Optional[bool] = None
  • SeqOptPlot's summary described two of its five methods

Verification

  • Code is untouched. For every changed file, the docstring-stripped AST is identical to master — checked per branch and again after the merge.
  • Docstring checker 0 defects; signature drift 0.
  • ~4500 tests across the four subpackages; 342 api_tests after the merge.
  • The Beta Features page still renders, which matters: _beta.py parses versionadded to build its "Added in" column.

Two targets were missed and are reported rather than padded: CPP.run reached 139 not ~100, and feature_map 216 not ~120. Both are now parameter-count-bound, not prose-bound — feature_map documents 53 parameters at 3.1 lines each, which is pandas density. Going shorter would mean documenting real parameters in one line or not at all.

Addresses #581.

🤖 Generated with Claude Code

breimanntools and others added 9 commits September 18, 2026 20:16
The v1.1.0 preprocessor docstrings had grown into mechanism dumps that
buried what the methods actually deliver. This cuts the machinery and
leads with the biology, without touching a line of code.

* Drop the 75-row normalization / de-normalization table from
  StructurePreprocessor.__init__ (110-line Notes block down to 20). The
  recipes live in the feature registry; the docstring now states the
  contract, that values are in [0, 1] with NaN for unresolved positions.
* Slim the fetch_embeddings model table from six columns to four and
  replace the UniProt precomputed-embeddings essay with two sentences.
* Replace the std-aware clustering derivation in
  EmbeddingPreprocessor.build_cat with a plain-language summary, keeping
  both citations.
* Rewrite the pad_parts "Application" block as flowing prose, per the
  style guide's rule against bold rhetorical labels.
* Drop the AlphaFold URL-versioning and encode_domains no-bundled-runtime
  rationale, and cut the raw GitHub URLs that citations already cover.

Version directives now follow scikit-learn placement: the load_dataset
class-level versionchanged is gone (the release notes carry the history)
and its one still-relevant fact sits under the `verbose` parameter as a
one-line versionadded. Every symbol keeps exactly one versionadded.

load_dataset(random=True) is unseeded, so the docstring now says so
rather than leaving reproducibility implied.

Every .. include:: examples line, [Key]_ citation, See Also entry and
numpydoc section is preserved; the ASTs are identical once docstrings
are blanked.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
The design, sequence-analysis and pipeline docstrings had grown a layer of
implementation exposition on top of the original writing. This removes that
layer and lets the biology lead. Docstring text only; no signature, default,
parameter or logic change.

- Cut the mechanism dumps: the DEAP parity table from the SeqOpt class
  docstring (a dev/test-only oracle, the shipped runtime never imports DEAP),
  the stage-by-stage grid exposition from ap.find_features, and the
  predictor-build and render-dispatch detail from CPPStructurePlot.explore.
  Each now says what the result means for a protein and what the user must
  know to drive it.
- Drop both `.. versionchanged::` directives (SeqOpt `mode` default and the
  `constraints` object form). Both describe the current documented behaviour,
  which the parameter entry already states; the release notes carry the history.
- Collapse repeated version notes to one per symbol: 13 method-level
  `.. versionadded::` that only restated their own class's version
  (AAWindowSampler, CPPStructurePlot, DesignConstraints), and four
  parameter-level ones on `predict_samples`, which becomes a single
  symbol-level note matching its sibling pipelines.
- Tighten the repeated `constraints` and `lineage` parameter blocks.

Kept verbatim: every `**Experimental.**` marker, all `.. include::` example
links, all `See Also` entries, and all citations (verified per file).

Gates: check_docstrings 0 defects (11 advisory, unchanged from baseline),
doc_signature_drift 0, 1904 unit tests pass across protein_engineering,
design_constraints, seqopt, api, seq_analysis, aa_window_sampler, pipe and
cpp_structure_plot. An AST comparison with docstrings stripped confirms the
code is identical.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
The prediction subpackage read as engineering notes rather than documentation:
long mechanism restatements, version history repeated on every symbol, and
em dashes used as the default connector.

- Drop every class-level and method-level ``.. versionchanged::`` (10 of 11);
  the release notes carry the history. The one behaviour a user can still get
  wrong, ``ReliabilityModel.fit(ci=...)`` taking a fraction instead of a
  percent, stays as a one-line note under that parameter, as do the ``mcc``
  metric and the ``'dendrogram'`` kind.
- Keep exactly one ``.. versionadded::`` per public symbol, plus parameter- and
  attribute-level ones for members that arrived after their class
  (``ad_borderline``, ``ad_threshold_``, ``ad_method_``, ``label``,
  ``use_calibrated``, ``add_metrics``, ``X_ref``). Nine parameter-level
  directives that merely repeated their own method's 1.1.0 are gone.
- Cut the mechanism dumps: the trust list-table, the per-kind "Uses <param>,
  <param>" routing lists, the learning-curve fold arithmetic, the validation
  rules already enforced by the checks, and the "attributes carry a trailing
  underscore" notes. One plain sentence is kept wherever the behaviour would
  surprise (untouched test fold, calibration built from the first ensemble
  member, ``score_std`` meaning something else in ``AAPred.eval``).
- Lead with meaning: ``bind_groups`` now opens on the dependent samples and the
  inflated score, with the splitter mechanics following; ``predict_candidates``,
  ``predict_oof``, ``eval_selective`` and ``learning_curve`` likewise state the
  question before the machinery.
- Replace the em dashes of the touched docstrings with colons and commas, per
  the prose rules of the docstring style guide.

Docstrings only: the docstring-stripped AST of all nine files is unchanged.
The ``**Experimental.**`` markers, numpydoc sections, named Returns, Examples
includes, citations, and See Also entries are untouched.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
…tacks

The feature-engineering docstrings had absorbed the rationale of each
feature as it was added, burying what a reader needs at the point of use
under release history and mechanism prose.

- Remove all 15 class- and method-level ``versionchanged`` directives.
  Four that recorded a new parameter move under that parameter as a
  one-line ``versionadded`` (scikit-learn placement): ``tmd_len`` and
  ``strategy``/``batch``/``df_seq``/``df_parts_kws`` in SequenceFeature.
  The rest are dropped; the release notes carry the history.
- Cut the mechanism dumps: the bootstrap "Choosing the settings" and
  cross-cutting-wrapper prose, the batching memory/complexity paragraphs,
  the splits auto-cap explanation (one plain sentence now sits under
  ``split_kws``, where the surprise is), the long multi-class note, the
  ``_kws`` preamble, and the O(...) complexity notes on the label helpers.
- Replace the 13-item ``df_feat`` column enumeration in ``CPP.run`` with a
  pointer to the generated, test-guarded df_feat contract page.
- Drop duplicated text: ``simplify``'s ``strategy`` was documented twice,
  and CPPGrid's ``backend`` note restated its own parameter.

Docstrings only: the code AST is byte-identical once docstrings are
stripped. Biology-facing explanations, ``**Experimental.**`` markers,
numpydoc structure, named Returns, examples includes, citations and
See Also entries are preserved.

CPP.run 215 -> 139 lines, CPP.__init__ 118 -> 79, CPP.run_num 177 -> 141,
CPPPlot.feature_map 240 -> 216; 19 docstrings, 1987 -> 1759 lines total.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
68 directives named a pre-release version - 0.1.0, 0.1.2, 0.1.3 - on the
founding API: CPP, AAclust, SequenceFeature, load_dataset, dPULearn,
TreeModel, the plotting functions. "Added in version 0.1.0" names a
version no user installed, and from the reader's point of view those
symbols have been there since the beginning.

They now read 1.0.0, the first public release. Every public symbol still
carries exactly one stamp, so the scikit-learn convention and the
checker's NO-VERSIONADDED rule are both untouched; only the floor moved.

This is a deliberate convention rather than a literal record: those
symbols did exist in the 0.1.x tags. The claim being made is "present
since the first release", which is what a reader who found the package on
PyPI understands, and the guide now says so.

Docstring text only: the docstring-stripped AST of all 21 files is
identical to origin/master. 342 api_tests pass and the docstring checker
reports 0 defects.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
@codecov

codecov Bot commented Sep 18, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 95.39%. Comparing base (fc8e21e) to head (61accd4).

Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff           @@
##           master     #589   +/-   ##
=======================================
  Coverage   95.39%   95.39%           
=======================================
  Files         222      222           
  Lines       23387    23387           
  Branches     4073     4073           
=======================================
  Hits        22309    22309           
  Misses        631      631           
  Partials      447      447           
Files with missing lines Coverage Δ
aaanalysis/data_handling/_embed_preproc.py 100.00% <ø> (ø)
aaanalysis/data_handling/_load_dataset.py 100.00% <ø> (ø)
aaanalysis/data_handling/_load_features.py 100.00% <ø> (ø)
aaanalysis/data_handling/_load_scales.py 95.77% <ø> (ø)
aaanalysis/data_handling/_seq_preproc.py 100.00% <ø> (ø)
aaanalysis/data_handling_pro/_annot_preproc.py 99.20% <ø> (ø)
aaanalysis/data_handling_pro/_struct_preproc.py 96.14% <ø> (ø)
aaanalysis/explainable_ai/_tree_model.py 95.36% <ø> (ø)
aaanalysis/explainable_ai_pro/_shap_model.py 98.41% <ø> (ø)
aaanalysis/feature_engineering/_aaclust.py 97.55% <ø> (ø)
... and 31 more
Components Coverage Δ
cpp_core 95.97% <ø> (ø)
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants