From 9b090d665102c89c2f4d036d5a286c451907096e Mon Sep 17 00:00:00 2001 From: Stephan Breimann Date: Fri, 18 Sep 2026 19:34:03 +0200 Subject: [PATCH] docs: TMD means Target Middle Domain, and it is already the abstraction The published Feature Identification chapter has always said "Target Middle Domain (TMD)". The glossaries had drifted away from it: CONTEXT.md called the part vocabulary "TMD-centric" and "semantically wrong" for non-domain levels and told the reader to "replace the generic tmd with the specific domain name", and docs/source/index/glossary.rst said the same. Read literally, that describes a defect, so it kept generating proposals to add a region-naming layer to CPP. Three independent planning passes - a codex planner, a subagent and this session - each reached that same wrong conclusion from the text alone, and each designed a region= parameter, a rename mapping and a place to store it. There is nothing to rename. The vocabulary is a geometry: one target middle domain flanked by two juxta middle domains, which applies unchanged to a Pfam domain, a cleavage-site window or a whole chain. A feature id reading TMD-Segment(2,4)-ANDN920101 does not assert a transmembrane helix. So the statement now appears in six places, all pointing at one canon: - feature_identification.rst gains an `important` admonition and the `part_vocabulary` anchor (and a JMD-C typo fix: it read "N-terminal") - glossary.rst: the part and df_parts entries, cross-referencing the anchor - CONTEXT.md: the part entry, with a note that it had drifted - _schemas.py: the df_parts description, rendered into the schema page - _constants.py: the comment above DICT_PART_LABEL, replacing the stale "reserved for the #27 region abstraction" note - CLAUDE.md: an always-loaded rule, so an agent reads it before designing Comment and description strings only in the two Python files; no code changed, and CPP is untouched. df_schemas.rst is regenerated from the edited schema so its sync test passes. Co-Authored-By: Claude Fable 5.1 --- CLAUDE.md | 10 +++++++ CONTEXT.md | 26 ++++++++++++++----- aaanalysis/_constants.py | 4 ++- aaanalysis/_schemas.py | 6 ++++- docs/source/index/glossary.rst | 14 ++++++---- .../index/usage_principles/df_schemas.rst | 2 +- .../feature_identification.rst | 14 +++++++++- 7 files changed, 60 insertions(+), 16 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index f12751629..32bb5b2bf 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -105,6 +105,16 @@ The detailed rules live in the path-scoped files above; these are the highest-risk traps worth keeping front-of-mind every session. - **No `print(...)` in library code.** Use `ut.print_out(...)`. +- **`tmd` means TARGET MIDDLE DOMAIN — it is already the general abstraction.** The part + vocabulary (`tmd` + the `jmd_n` / `jmd_c` flanks and their composites) is a *geometry*, + not a membrane claim: the same names serve a Pfam domain, a cleavage-site window or a + whole chain, and `TMD-Segment(2,4)-ANDN920101` does **not** assert a transmembrane helix. + So **never propose "generalizing" part naming** — no `region=` parameter, no per-domain + rename layer, no mapping carried in `df.attrs`. Three independent planning passes each + proposed exactly that, purely because the glossary once called the vocabulary + "semantically wrong"; the text is fixed, and the rule is: renaming parts is cosmetic and + CPP is not changed for cosmetics. Canon: the *Feature Identification* chapter + (`part_vocabulary` anchor), the `part` entry in `CONTEXT.md`, and `docs/source/index/glossary.rst`. - **No ADR references in project code or GitHub (hard rule).** Never cite an ADR (`ADR-0001`, "see ADR-0001", a `docs/adr/...` path) in **any `.py` file** (docstrings **and** `#` comments, library **and** test) or in **GitHub** diff --git a/CONTEXT.md b/CONTEXT.md index c90ec922e..f667de62d 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -27,15 +27,27 @@ The single rule governing `tmd_start` / `tmd_stop` (`ut.COL_TMD_START` / `ut.COL _Avoid_: 0-based, half-open / exclusive-stop, `len()`-style stop (a stop equal to `start + length` reads as exclusive — it is not). **part**: -A named region of a protein over which a **split** operates and a scale is averaged; the `PART` field of a feature id (`PART-SPLIT-SCALE`). Parts are the columns of `df_parts`, produced by `SequenceFeature.get_df_parts`. The default vocabulary is **TMD-centric** — `jmd_n` / `tmd` / `jmd_c` (plus composites like `jmd_n_tmd_n`) — which fits **domain-level** tasks but is *semantically wrong* for other levels, so part naming should follow the **prediction level**: -- **Domain level:** replace the generic `tmd` with the **specific domain name** where known (e.g. the Pfam / InterPro domain), rather than the placeholder "tmd". -- **Residue level (cleavage / between-residues):** name positions by the **Schechter–Berger** convention — `… P2 · P1 │ P1′ · P2′ …` around the scissile bond (`│` = cleavage site; see **P1 anchor / source position**), not "tmd". -- **Protein level:** the whole chain is a single part; use a neutral name (e.g. `seq`, or N-term / core / C-term thirds), not "tmd". -First-class user-defined / renamed regions are tracked by **#27** (region abstraction); today a part is chosen from the predefined family. -_Avoid_: region (reserved for the #27 abstraction), domain (a part may be a window or sub-region, not a whole domain), segment (a split type). +A named region of a protein over which a **split** operates and a scale is averaged; the `PART` field of a feature id (`PART-SPLIT-SCALE`). Parts are the columns of `df_parts`, produced by `SequenceFeature.get_df_parts`. + +> **`tmd` means TARGET MIDDLE DOMAIN. It is already the general abstraction — do not "generalize" it.** +> +> The part vocabulary is a **geometry**, not a biological claim: one **target middle domain** (`tmd`) flanked by two **juxta middle domains** (`jmd_n`, `jmd_c`), plus their composites (`jmd_n_tmd_n`, `tmd_c_jmd_c`, …). The letters are historical — the abstraction is "the span of interest and its two flanks", and it applies unchanged to a Pfam domain, a kinase domain, a cleavage-site window or a whole chain. A feature id reading `TMD-Segment(2,4)-ANDN920101` is therefore **not** claiming a transmembrane helix; it names the target span. +> +> The published statement of this is the *Feature Identification* chapter of the docs +> (`docs/source/index/usage_principles/feature_identification.rst`, anchor `part_vocabulary`), +> which has said "Target Middle Domain" all along; this glossary had drifted away from it. +> +> This paragraph exists because the earlier wording here (the vocabulary is "TMD-centric", "semantically wrong" for non-membrane levels, "replace the generic `tmd` with the specific domain name") read as a defect and repeatedly generated proposals to add a region-naming layer to CPP. Three independent planning passes reached that same wrong conclusion from this text alone. There is nothing to rename: **renaming parts is cosmetic, and CPP is not to be changed for it.** + +How the geometry maps onto each **prediction level** (the *names* stay the same; only what the span means changes): +- **Domain level:** `tmd` is the domain of interest (e.g. a Pfam / InterPro domain), `jmd_n` / `jmd_c` its flanks. +- **Residue level (cleavage / between-residues):** `tmd` is the window around the scissile bond. When describing positions in prose, the **Schechter–Berger** convention (`… P2 · P1 │ P1′ · P2′ …`, `│` = cleavage site; see **P1 anchor / source position**) is the reader-facing vocabulary — it is a way of *talking about positions*, not a different part vocabulary. +- **Protein level:** the whole chain is the target span. + +_Avoid_: region (an informal synonym; **part** is the term), domain (a part may be a window or sub-region, not a whole domain), segment (a split type). **part label**: -The fixed, human-readable display string for a `part` token (e.g. `tmd` → "TMD", `jmd_n_tmd_n` → "JMD-N+TMD-N"), defined once in `ut.DICT_PART_LABEL` and used when rendering a feature id as prose. Deliberately *not* called a "region" (that noun is reserved for #27); a part label is purely cosmetic and changes no behavior. +The fixed, human-readable display string for a `part` token (e.g. `tmd` → "TMD", `jmd_n_tmd_n` → "JMD-N+TMD-N"), defined once in `ut.DICT_PART_LABEL` and used when rendering a feature id as prose. The term is **part label**, not "region"; a part label is purely cosmetic and changes no behavior. **feature description**: One standardized, human-readable sentence built deterministically from a `PART-SPLIT-SCALE` feature id by `SequenceFeature.get_feature_descriptions`: it joins the **part label**, the **split** rendered as a phrase (e.g. "segment 2 of 4"), the residue positions, and the scale's AAontology name / category / subcategory from `df_cat`. Additive only — the `feature` id string is unchanged — and optionally carried as the `feature_description` (`ut.COL_FEAT_DES`) column of `df_feat`. Distinct from the compact **feature name** (`get_feature_names`, `'subcategory [positions]'`), which drops the part and category. diff --git a/aaanalysis/_constants.py b/aaanalysis/_constants.py index 43563c925..ef15c541f 100644 --- a/aaanalysis/_constants.py +++ b/aaanalysis/_constants.py @@ -45,7 +45,9 @@ def _folder_path(super_folder, folder_name): # Canonical, human-readable label per sequence part (the PART field of a # PART-SPLIT-SCALE feature id). Single source of the part-label vocabulary used by # SequenceFeature.get_feature_descriptions; keys cover every part in LIST_ALL_PARTS. -# ('region' is deliberately avoided here — reserved for the #27 region abstraction.) +# 'tmd' is the TARGET MIDDLE DOMAIN: the span of interest, flanked by jmd_n / jmd_c. +# The vocabulary is a geometry, not a membrane claim, so it is already general — do not +# add a per-domain renaming layer on top of it (see the 'part' entry in CONTEXT.md). DICT_PART_LABEL = {"tmd": "TMD", "tmd_e": "extended TMD", "tmd_n": "TMD-N", diff --git a/aaanalysis/_schemas.py b/aaanalysis/_schemas.py index e36b687d0..3ad1b160a 100644 --- a/aaanalysis/_schemas.py +++ b/aaanalysis/_schemas.py @@ -133,7 +133,11 @@ def _field(dtype, description, *, required=True, nullable=False, unique=False, "description": ( "Sequence parts table consumed by CPP / SequenceFeature; one row per " "sequence. Columns are DYNAMIC: one column per selected sequence part, named " - "by the part vocabulary; each value is the part's amino acid subsequence."), + "by the part vocabulary; each value is the part's amino acid subsequence. " + "In that vocabulary 'tmd' is the TARGET MIDDLE DOMAIN (the target span of " + "interest) and 'jmd_n' / 'jmd_c' are its flanks: a geometry, not a membrane " + "claim, which is why the same names serve a domain, a cleavage-site window or " + "a whole chain. See the Feature Identification chapter of the documentation."), "dynamic_columns": { "name_from": "LIST_ALL_PARTS", "allowed_names": list(LIST_ALL_PARTS), diff --git a/docs/source/index/glossary.rst b/docs/source/index/glossary.rst index 8576cdd85..513bfd3bf 100644 --- a/docs/source/index/glossary.rst +++ b/docs/source/index/glossary.rst @@ -23,8 +23,9 @@ Sequences & data objects ``tmd_start`` / ``tmd_stop``. df_parts - Wide table with one column per :term:`part` (``tmd``, ``jmd_n``, - ``jmd_c``, …), produced by :meth:`~aaanalysis.SequenceFeature.get_df_parts`. + Wide table with one column per :term:`part` (``tmd`` = Target Middle + Domain, ``jmd_n``, ``jmd_c``, …), produced by + :meth:`~aaanalysis.SequenceFeature.get_df_parts`. df_feat Ranked feature table: ``feature`` id, ``abs_auc``, ``mean_dif``, @@ -36,9 +37,12 @@ Sequences & data objects part A named region of a sequence over which a :term:`split` operates and a :term:`scale` is averaged; the ``PART`` field of a feature id - (``PART-SPLIT-SCALE``). The default vocabulary is TMD-centric (``jmd_n`` / - ``tmd`` / ``jmd_c`` and composites); name parts after the - :term:`prediction level` when that fits better. + (``PART-SPLIT-SCALE``). The vocabulary is ``tmd`` / ``jmd_n`` / ``jmd_c`` + and their composites, where **TMD is the Target Middle Domain** — the + target span of interest — and the JMDs are its two flanks. It is a + **geometry, not a biological claim**: the same names serve a domain, a + cleavage-site window or a whole chain, so parts are not renamed per data + set. See :ref:`part_vocabulary`. scale A mapping from each amino acid to a real number — a physicochemical diff --git a/docs/source/index/usage_principles/df_schemas.rst b/docs/source/index/usage_principles/df_schemas.rst index e0c7a8914..dacacb482 100644 --- a/docs/source/index/usage_principles/df_schemas.rst +++ b/docs/source/index/usage_principles/df_schemas.rst @@ -104,7 +104,7 @@ Accepted column formats (besides the always-required columns): ``df_parts`` ------------ -Sequence parts table consumed by CPP / SequenceFeature; one row per sequence. Columns are DYNAMIC: one column per selected sequence part, named by the part vocabulary; each value is the part's amino acid subsequence. +Sequence parts table consumed by CPP / SequenceFeature; one row per sequence. Columns are DYNAMIC: one column per selected sequence part, named by the part vocabulary; each value is the part's amino acid subsequence. In that vocabulary 'tmd' is the TARGET MIDDLE DOMAIN (the target span of interest) and 'jmd_n' / 'jmd_c' are its flanks: a geometry, not a membrane claim, which is why the same names serve a domain, a cleavage-site window or a whole chain. See the Feature Identification chapter of the documentation. Dynamic columns (dtype: str, nullable: False): Amino acid subsequence of the named sequence part. Column names are drawn from the part vocabulary: ``tmd``, ``tmd_e``, ``tmd_n``, ``tmd_c``, ``jmd_n``, ``jmd_c``, ``ext_c``, ``ext_n``, ``tmd_jmd``, ``jmd_n_tmd_n``, ``tmd_c_jmd_c``, ``ext_n_tmd_n``, ``tmd_c_ext_c``. diff --git a/docs/source/index/usage_principles/feature_identification.rst b/docs/source/index/usage_principles/feature_identification.rst index fcd0d4803..416e96247 100755 --- a/docs/source/index/usage_principles/feature_identification.rst +++ b/docs/source/index/usage_principles/feature_identification.rst @@ -25,8 +25,20 @@ The core idea of CPP is its feature concept: Scheme of CPP feature (**Part-Split-Scale** combination) with example of feature creation, from [Breimann25]_. +.. _part_vocabulary: + All possible parts are sub-parts or combinations of the **Target Middle Domain (TMD)**, -**Juxta Middle Domain N-terminal (JMD-N)**, and **Juxta Middle Domain N-terminal (JMD-C)**. +**Juxta Middle Domain N-terminal (JMD-N)**, and **Juxta Middle Domain C-terminal (JMD-C)**. + +.. admonition:: TMD means *Target Middle Domain*, not *transmembrane domain* + :class: important + + The part vocabulary is a **geometry**, not a biological claim: one target span (``tmd``) + flanked by two juxta spans (``jmd_n``, ``jmd_c``), plus their composites. It is already + general and applies unchanged to a Pfam domain, a kinase domain, a cleavage-site window or + a whole chain. A feature id such as ``TMD-Segment(2,4)-ANDN920101`` therefore does **not** + assert a transmembrane helix; it names the target span. The letters are historical, as the + paragraph below explains, and there is nothing to rename per data set. .. figure:: /_artwork/schemes/scheme_CPP2.png