Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,6 +105,16 @@ The detailed rules live in the path-scoped files above; these are the
highest-risk traps worth keeping front-of-mind every session.

- **No `print(...)` in library code.** Use `ut.print_out(...)`.
- **`tmd` means TARGET MIDDLE DOMAIN — it is already the general abstraction.** The part
vocabulary (`tmd` + the `jmd_n` / `jmd_c` flanks and their composites) is a *geometry*,
not a membrane claim: the same names serve a Pfam domain, a cleavage-site window or a
whole chain, and `TMD-Segment(2,4)-ANDN920101` does **not** assert a transmembrane helix.
So **never propose "generalizing" part naming** — no `region=` parameter, no per-domain
rename layer, no mapping carried in `df.attrs`. Three independent planning passes each
proposed exactly that, purely because the glossary once called the vocabulary
"semantically wrong"; the text is fixed, and the rule is: renaming parts is cosmetic and
CPP is not changed for cosmetics. Canon: the *Feature Identification* chapter
(`part_vocabulary` anchor), the `part` entry in `CONTEXT.md`, and `docs/source/index/glossary.rst`.
- **No ADR references in project code or GitHub (hard rule).** Never cite an
ADR (`ADR-0001`, "see ADR-0001", a `docs/adr/...` path) in **any `.py` file**
(docstrings **and** `#` comments, library **and** test) or in **GitHub**
Expand Down
26 changes: 19 additions & 7 deletions CONTEXT.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,15 +27,27 @@ The single rule governing `tmd_start` / `tmd_stop` (`ut.COL_TMD_START` / `ut.COL
_Avoid_: 0-based, half-open / exclusive-stop, `len()`-style stop (a stop equal to `start + length` reads as exclusive — it is not).

**part**:
A named region of a protein over which a **split** operates and a scale is averaged; the `PART` field of a feature id (`PART-SPLIT-SCALE`). Parts are the columns of `df_parts`, produced by `SequenceFeature.get_df_parts`. The default vocabulary is **TMD-centric** — `jmd_n` / `tmd` / `jmd_c` (plus composites like `jmd_n_tmd_n`) — which fits **domain-level** tasks but is *semantically wrong* for other levels, so part naming should follow the **prediction level**:
- **Domain level:** replace the generic `tmd` with the **specific domain name** where known (e.g. the Pfam / InterPro domain), rather than the placeholder "tmd".
- **Residue level (cleavage / between-residues):** name positions by the **Schechter–Berger** convention — `… P2 · P1 │ P1′ · P2′ …` around the scissile bond (`│` = cleavage site; see **P1 anchor / source position**), not "tmd".
- **Protein level:** the whole chain is a single part; use a neutral name (e.g. `seq`, or N-term / core / C-term thirds), not "tmd".
First-class user-defined / renamed regions are tracked by **#27** (region abstraction); today a part is chosen from the predefined family.
_Avoid_: region (reserved for the #27 abstraction), domain (a part may be a window or sub-region, not a whole domain), segment (a split type).
A named region of a protein over which a **split** operates and a scale is averaged; the `PART` field of a feature id (`PART-SPLIT-SCALE`). Parts are the columns of `df_parts`, produced by `SequenceFeature.get_df_parts`.

> **`tmd` means TARGET MIDDLE DOMAIN. It is already the general abstraction — do not "generalize" it.**
>
> The part vocabulary is a **geometry**, not a biological claim: one **target middle domain** (`tmd`) flanked by two **juxta middle domains** (`jmd_n`, `jmd_c`), plus their composites (`jmd_n_tmd_n`, `tmd_c_jmd_c`, …). The letters are historical — the abstraction is "the span of interest and its two flanks", and it applies unchanged to a Pfam domain, a kinase domain, a cleavage-site window or a whole chain. A feature id reading `TMD-Segment(2,4)-ANDN920101` is therefore **not** claiming a transmembrane helix; it names the target span.
>
> The published statement of this is the *Feature Identification* chapter of the docs
> (`docs/source/index/usage_principles/feature_identification.rst`, anchor `part_vocabulary`),
> which has said "Target Middle Domain" all along; this glossary had drifted away from it.
>
> This paragraph exists because the earlier wording here (the vocabulary is "TMD-centric", "semantically wrong" for non-membrane levels, "replace the generic `tmd` with the specific domain name") read as a defect and repeatedly generated proposals to add a region-naming layer to CPP. Three independent planning passes reached that same wrong conclusion from this text alone. There is nothing to rename: **renaming parts is cosmetic, and CPP is not to be changed for it.**

How the geometry maps onto each **prediction level** (the *names* stay the same; only what the span means changes):
- **Domain level:** `tmd` is the domain of interest (e.g. a Pfam / InterPro domain), `jmd_n` / `jmd_c` its flanks.
- **Residue level (cleavage / between-residues):** `tmd` is the window around the scissile bond. When describing positions in prose, the **Schechter–Berger** convention (`… P2 · P1 │ P1′ · P2′ …`, `│` = cleavage site; see **P1 anchor / source position**) is the reader-facing vocabulary — it is a way of *talking about positions*, not a different part vocabulary.
- **Protein level:** the whole chain is the target span.

_Avoid_: region (an informal synonym; **part** is the term), domain (a part may be a window or sub-region, not a whole domain), segment (a split type).

**part label**:
The fixed, human-readable display string for a `part` token (e.g. `tmd` → "TMD", `jmd_n_tmd_n` → "JMD-N+TMD-N"), defined once in `ut.DICT_PART_LABEL` and used when rendering a feature id as prose. Deliberately *not* called a "region" (that noun is reserved for #27); a part label is purely cosmetic and changes no behavior.
The fixed, human-readable display string for a `part` token (e.g. `tmd` → "TMD", `jmd_n_tmd_n` → "JMD-N+TMD-N"), defined once in `ut.DICT_PART_LABEL` and used when rendering a feature id as prose. The term is **part label**, not "region"; a part label is purely cosmetic and changes no behavior.

**feature description**:
One standardized, human-readable sentence built deterministically from a `PART-SPLIT-SCALE` feature id by `SequenceFeature.get_feature_descriptions`: it joins the **part label**, the **split** rendered as a phrase (e.g. "segment 2 of 4"), the residue positions, and the scale's AAontology name / category / subcategory from `df_cat`. Additive only — the `feature` id string is unchanged — and optionally carried as the `feature_description` (`ut.COL_FEAT_DES`) column of `df_feat`. Distinct from the compact **feature name** (`get_feature_names`, `'subcategory [positions]'`), which drops the part and category.
Expand Down
4 changes: 3 additions & 1 deletion aaanalysis/_constants.py
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,9 @@ def _folder_path(super_folder, folder_name):
# Canonical, human-readable label per sequence part (the PART field of a
# PART-SPLIT-SCALE feature id). Single source of the part-label vocabulary used by
# SequenceFeature.get_feature_descriptions; keys cover every part in LIST_ALL_PARTS.
# ('region' is deliberately avoided here — reserved for the #27 region abstraction.)
# 'tmd' is the TARGET MIDDLE DOMAIN: the span of interest, flanked by jmd_n / jmd_c.
# The vocabulary is a geometry, not a membrane claim, so it is already general — do not
# add a per-domain renaming layer on top of it (see the 'part' entry in CONTEXT.md).
DICT_PART_LABEL = {"tmd": "TMD",
"tmd_e": "extended TMD",
"tmd_n": "TMD-N",
Expand Down
6 changes: 5 additions & 1 deletion aaanalysis/_schemas.py
Original file line number Diff line number Diff line change
Expand Up @@ -133,7 +133,11 @@ def _field(dtype, description, *, required=True, nullable=False, unique=False,
"description": (
"Sequence parts table consumed by CPP / SequenceFeature; one row per "
"sequence. Columns are DYNAMIC: one column per selected sequence part, named "
"by the part vocabulary; each value is the part's amino acid subsequence."),
"by the part vocabulary; each value is the part's amino acid subsequence. "
"In that vocabulary 'tmd' is the TARGET MIDDLE DOMAIN (the target span of "
"interest) and 'jmd_n' / 'jmd_c' are its flanks: a geometry, not a membrane "
"claim, which is why the same names serve a domain, a cleavage-site window or "
"a whole chain. See the Feature Identification chapter of the documentation."),
"dynamic_columns": {
"name_from": "LIST_ALL_PARTS",
"allowed_names": list(LIST_ALL_PARTS),
Expand Down
14 changes: 9 additions & 5 deletions docs/source/index/glossary.rst
Original file line number Diff line number Diff line change
Expand Up @@ -23,8 +23,9 @@ Sequences & data objects
``tmd_start`` / ``tmd_stop``.

df_parts
Wide table with one column per :term:`part` (``tmd``, ``jmd_n``,
``jmd_c``, …), produced by :meth:`~aaanalysis.SequenceFeature.get_df_parts`.
Wide table with one column per :term:`part` (``tmd`` = Target Middle
Domain, ``jmd_n``, ``jmd_c``, …), produced by
:meth:`~aaanalysis.SequenceFeature.get_df_parts`.

df_feat
Ranked feature table: ``feature`` id, ``abs_auc``, ``mean_dif``,
Expand All @@ -36,9 +37,12 @@ Sequences & data objects
part
A named region of a sequence over which a :term:`split` operates and a
:term:`scale` is averaged; the ``PART`` field of a feature id
(``PART-SPLIT-SCALE``). The default vocabulary is TMD-centric (``jmd_n`` /
``tmd`` / ``jmd_c`` and composites); name parts after the
:term:`prediction level` when that fits better.
(``PART-SPLIT-SCALE``). The vocabulary is ``tmd`` / ``jmd_n`` / ``jmd_c``
and their composites, where **TMD is the Target Middle Domain** — the
target span of interest — and the JMDs are its two flanks. It is a
**geometry, not a biological claim**: the same names serve a domain, a
cleavage-site window or a whole chain, so parts are not renamed per data
set. See :ref:`part_vocabulary`.

scale
A mapping from each amino acid to a real number — a physicochemical
Expand Down
2 changes: 1 addition & 1 deletion docs/source/index/usage_principles/df_schemas.rst
Original file line number Diff line number Diff line change
Expand Up @@ -104,7 +104,7 @@ Accepted column formats (besides the always-required columns):
``df_parts``
------------

Sequence parts table consumed by CPP / SequenceFeature; one row per sequence. Columns are DYNAMIC: one column per selected sequence part, named by the part vocabulary; each value is the part's amino acid subsequence.
Sequence parts table consumed by CPP / SequenceFeature; one row per sequence. Columns are DYNAMIC: one column per selected sequence part, named by the part vocabulary; each value is the part's amino acid subsequence. In that vocabulary 'tmd' is the TARGET MIDDLE DOMAIN (the target span of interest) and 'jmd_n' / 'jmd_c' are its flanks: a geometry, not a membrane claim, which is why the same names serve a domain, a cleavage-site window or a whole chain. See the Feature Identification chapter of the documentation.

Dynamic columns (dtype: str, nullable: False): Amino acid subsequence of the named sequence part. Column names are drawn from the part vocabulary: ``tmd``, ``tmd_e``, ``tmd_n``, ``tmd_c``, ``jmd_n``, ``jmd_c``, ``ext_c``, ``ext_n``, ``tmd_jmd``, ``jmd_n_tmd_n``, ``tmd_c_jmd_c``, ``ext_n_tmd_n``, ``tmd_c_ext_c``.

Expand Down
14 changes: 13 additions & 1 deletion docs/source/index/usage_principles/feature_identification.rst
Original file line number Diff line number Diff line change
Expand Up @@ -25,8 +25,20 @@ The core idea of CPP is its feature concept:

Scheme of CPP feature (**Part-Split-Scale** combination) with example of feature creation, from [Breimann25]_.

.. _part_vocabulary:

All possible parts are sub-parts or combinations of the **Target Middle Domain (TMD)**,
**Juxta Middle Domain N-terminal (JMD-N)**, and **Juxta Middle Domain N-terminal (JMD-C)**.
**Juxta Middle Domain N-terminal (JMD-N)**, and **Juxta Middle Domain C-terminal (JMD-C)**.

.. admonition:: TMD means *Target Middle Domain*, not *transmembrane domain*
:class: important

The part vocabulary is a **geometry**, not a biological claim: one target span (``tmd``)
flanked by two juxta spans (``jmd_n``, ``jmd_c``), plus their composites. It is already
general and applies unchanged to a Pfam domain, a kinase domain, a cleavage-site window or
a whole chain. A feature id such as ``TMD-Segment(2,4)-ANDN920101`` therefore does **not**
assert a transmembrane helix; it names the target span. The letters are historical, as the
paragraph below explains, and there is nothing to rename per data set.

.. figure:: /_artwork/schemes/scheme_CPP2.png

Expand Down
Loading