docs: compact the docstrings and fix the version stamps - #589
Open
breimanntools wants to merge 9 commits into
Open
breimanntools wants to merge 9 commits into
breimanntools wants to merge 9 commits into
Conversation
The v1.1.0 preprocessor docstrings had grown into mechanism dumps that buried what the methods actually deliver. This cuts the machinery and leads with the biology, without touching a line of code. * Drop the 75-row normalization / de-normalization table from StructurePreprocessor.__init__ (110-line Notes block down to 20). The recipes live in the feature registry; the docstring now states the contract, that values are in [0, 1] with NaN for unresolved positions. * Slim the fetch_embeddings model table from six columns to four and replace the UniProt precomputed-embeddings essay with two sentences. * Replace the std-aware clustering derivation in EmbeddingPreprocessor.build_cat with a plain-language summary, keeping both citations. * Rewrite the pad_parts "Application" block as flowing prose, per the style guide's rule against bold rhetorical labels. * Drop the AlphaFold URL-versioning and encode_domains no-bundled-runtime rationale, and cut the raw GitHub URLs that citations already cover. Version directives now follow scikit-learn placement: the load_dataset class-level versionchanged is gone (the release notes carry the history) and its one still-relevant fact sits under the `verbose` parameter as a one-line versionadded. Every symbol keeps exactly one versionadded. load_dataset(random=True) is unseeded, so the docstring now says so rather than leaving reproducibility implied. Every .. include:: examples line, [Key]_ citation, See Also entry and numpydoc section is preserved; the ASTs are identical once docstrings are blanked. Co-Authored-By: Claude Fable 5.1 <[email protected]>
The design, sequence-analysis and pipeline docstrings had grown a layer of implementation exposition on top of the original writing. This removes that layer and lets the biology lead. Docstring text only; no signature, default, parameter or logic change. - Cut the mechanism dumps: the DEAP parity table from the SeqOpt class docstring (a dev/test-only oracle, the shipped runtime never imports DEAP), the stage-by-stage grid exposition from ap.find_features, and the predictor-build and render-dispatch detail from CPPStructurePlot.explore. Each now says what the result means for a protein and what the user must know to drive it. - Drop both `.. versionchanged::` directives (SeqOpt `mode` default and the `constraints` object form). Both describe the current documented behaviour, which the parameter entry already states; the release notes carry the history. - Collapse repeated version notes to one per symbol: 13 method-level `.. versionadded::` that only restated their own class's version (AAWindowSampler, CPPStructurePlot, DesignConstraints), and four parameter-level ones on `predict_samples`, which becomes a single symbol-level note matching its sibling pipelines. - Tighten the repeated `constraints` and `lineage` parameter blocks. Kept verbatim: every `**Experimental.**` marker, all `.. include::` example links, all `See Also` entries, and all citations (verified per file). Gates: check_docstrings 0 defects (11 advisory, unchanged from baseline), doc_signature_drift 0, 1904 unit tests pass across protein_engineering, design_constraints, seqopt, api, seq_analysis, aa_window_sampler, pipe and cpp_structure_plot. An AST comparison with docstrings stripped confirms the code is identical. Co-Authored-By: Claude Fable 5.1 <[email protected]>
The prediction subpackage read as engineering notes rather than documentation: long mechanism restatements, version history repeated on every symbol, and em dashes used as the default connector. - Drop every class-level and method-level ``.. versionchanged::`` (10 of 11); the release notes carry the history. The one behaviour a user can still get wrong, ``ReliabilityModel.fit(ci=...)`` taking a fraction instead of a percent, stays as a one-line note under that parameter, as do the ``mcc`` metric and the ``'dendrogram'`` kind. - Keep exactly one ``.. versionadded::`` per public symbol, plus parameter- and attribute-level ones for members that arrived after their class (``ad_borderline``, ``ad_threshold_``, ``ad_method_``, ``label``, ``use_calibrated``, ``add_metrics``, ``X_ref``). Nine parameter-level directives that merely repeated their own method's 1.1.0 are gone. - Cut the mechanism dumps: the trust list-table, the per-kind "Uses <param>, <param>" routing lists, the learning-curve fold arithmetic, the validation rules already enforced by the checks, and the "attributes carry a trailing underscore" notes. One plain sentence is kept wherever the behaviour would surprise (untouched test fold, calibration built from the first ensemble member, ``score_std`` meaning something else in ``AAPred.eval``). - Lead with meaning: ``bind_groups`` now opens on the dependent samples and the inflated score, with the splitter mechanics following; ``predict_candidates``, ``predict_oof``, ``eval_selective`` and ``learning_curve`` likewise state the question before the machinery. - Replace the em dashes of the touched docstrings with colons and commas, per the prose rules of the docstring style guide. Docstrings only: the docstring-stripped AST of all nine files is unchanged. The ``**Experimental.**`` markers, numpydoc sections, named Returns, Examples includes, citations, and See Also entries are untouched. Co-Authored-By: Claude Fable 5.1 <[email protected]>
…tacks The feature-engineering docstrings had absorbed the rationale of each feature as it was added, burying what a reader needs at the point of use under release history and mechanism prose. - Remove all 15 class- and method-level ``versionchanged`` directives. Four that recorded a new parameter move under that parameter as a one-line ``versionadded`` (scikit-learn placement): ``tmd_len`` and ``strategy``/``batch``/``df_seq``/``df_parts_kws`` in SequenceFeature. The rest are dropped; the release notes carry the history. - Cut the mechanism dumps: the bootstrap "Choosing the settings" and cross-cutting-wrapper prose, the batching memory/complexity paragraphs, the splits auto-cap explanation (one plain sentence now sits under ``split_kws``, where the surprise is), the long multi-class note, the ``_kws`` preamble, and the O(...) complexity notes on the label helpers. - Replace the 13-item ``df_feat`` column enumeration in ``CPP.run`` with a pointer to the generated, test-guarded df_feat contract page. - Drop duplicated text: ``simplify``'s ``strategy`` was documented twice, and CPPGrid's ``backend`` note restated its own parameter. Docstrings only: the code AST is byte-identical once docstrings are stripped. Biology-facing explanations, ``**Experimental.**`` markers, numpydoc structure, named Returns, examples includes, citations and See Also entries are preserved. CPP.run 215 -> 139 lines, CPP.__init__ 118 -> 79, CPP.run_num 177 -> 141, CPPPlot.feature_map 240 -> 216; 19 docstrings, 1987 -> 1759 lines total. Co-Authored-By: Claude Fable 5.1 <[email protected]>
68 directives named a pre-release version - 0.1.0, 0.1.2, 0.1.3 - on the founding API: CPP, AAclust, SequenceFeature, load_dataset, dPULearn, TreeModel, the plotting functions. "Added in version 0.1.0" names a version no user installed, and from the reader's point of view those symbols have been there since the beginning. They now read 1.0.0, the first public release. Every public symbol still carries exactly one stamp, so the scikit-learn convention and the checker's NO-VERSIONADDED rule are both untouched; only the floor moved. This is a deliberate convention rather than a literal record: those symbols did exist in the 0.1.x tags. The claim being made is "present since the first release", which is what a reader who found the package on PyPI understands, and the guide now says so. Docstring text only: the docstring-stripped AST of all 21 files is identical to origin/master. 342 api_tests pass and the docstring checker reports 0 defects. Co-Authored-By: Claude Fable 5.1 <[email protected]>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #589 +/- ##
=======================================
Coverage 95.39% 95.39%
=======================================
Files 222 222
Lines 23387 23387
Branches 4073 4073
=======================================
Hits 22309 22309
Misses 631 631
Partials 447 447
🚀 New features to boost your workflow:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The docstrings had absorbed the rationale of every feature as it was added, and the API pages opened with a stack of "Changed in version …" blocks before the reader learned what a method does. This is the cleanup, across four subpackages plus the version stamps.
720 insertions, 1440 deletions — net −720 lines, all of it docstring text.
What changed
versionchangedversionaddednaming a pre-release (0.1.x)CPP.runCPP.__init__ReliabilityModelStructurePreprocessor.__init__SeqOpt(class)The 3 surviving
versionchangedare all1.2.0, one line each, sitting under the parameter they describe — scikit-learn's placement. They are kept because a user hitting the next release needs them:cichanged from a percent to a fraction, so someone passing90.0would otherwise silently get wrong behaviour.Version stamps keep the scikit-learn convention — every public class and function still carries exactly one
.. versionadded::, and the checker rule is untouched. Only the floor moved: 68 directives naming 0.1.0 / 0.1.2 / 0.1.3 now read1.0.0, since the public history starts at the first release and "added in 0.1.0" names a version no user installed. That is a deliberate convention, and the style guide and.claude/rules/docstrings.mdboth now say so.What was cut, and what was kept
Cut: a 75-row normalization table in
StructurePreprocessor.__init__that duplicatedfeature_registry.NORMALIZATION_RECIPES; a 13-row SeqOpt→DEAP mapping (DEAP is a dev/test-only parity oracle the runtime never imports);CPP.run's 13-itemdf_featcolumn list, now a pointer to the generated contract page that a test keeps in sync; theO(batch_size × part_length × n_scales)memory prose; the "Why these defaults?" bootstrap essays; the_kwspreamble.Kept: the biology, and every load-bearing trap reduced to one plain sentence — splits auto-capping, the raw-PLM normalization trap in
run_num, the sample-levelmean_diftrap, the CPP-inside-each-fold leakage warning. Plus every**Experimental.**marker (17 before, 17 after), every[Key]_citation, every.. include:: examples/*.rst, every namedReturns.Three docstrings that were factually wrong
AAPredclaimed it "intentionally does not perform hyperparameter optimization" whilefit(optimize_hyperparams=True)runsGridSearchCVCPPPlot.feature_map'sseq_char_fillcarried "Now defaults toTrue"; the signature isOptional[bool] = NoneSeqOptPlot's summary described two of its five methodsVerification
master— checked per branch and again after the merge.api_testsafter the merge._beta.pyparsesversionaddedto build its "Added in" column.Two targets were missed and are reported rather than padded:
CPP.runreached 139 not ~100, andfeature_map216 not ~120. Both are now parameter-count-bound, not prose-bound —feature_mapdocuments 53 parameters at 3.1 lines each, which is pandas density. Going shorter would mean documenting real parameters in one line or not at all.Addresses #581.
🤖 Generated with Claude Code