From 3daa4cf7c21313b3117923fe84ac0c22ab3ba561 Mon Sep 17 00:00:00 2001 From: Stephan Breimann Date: Fri, 18 Sep 2026 23:55:56 +0200 Subject: [PATCH] docs: record that TMD was Transmembrane Domain before it was Target Middle Domain The wording added earlier today said TMD means Target Middle Domain "not transmembrane domain". That flattens a design decision into a correction, and it is wrong in the one place it matters most: in the gamma-secretase work the span really is the transmembrane domain. The history is the explanation. CPP was first applied to gamma-secretase substrates, where the span of interest is the transmembrane domain and its flanks are the juxtamembrane domains. When the method was generalized beyond membrane proteins the acronyms were deliberately kept and re-read - TMD as Target Middle Domain, JMD as Juxta Middle Domain. Same letters, same geometry, wider meaning, nothing to rename per data set. Stated in the three places a reader meets it: - the Feature Identification chapter, whose admonition now leads with the origin and says a feature id describes a transmembrane helix only when the data are membrane proteins, as in the case study below it - the part entry in CONTEXT.md, the internal glossary that caused the region-renaming proposals in the first place - the gamma-secretase use case, right where the reader meets jmd_n / tmd / jmd_c, since that analysis is the one the vocabulary is named after and the literal reading is the correct one there Documentation only. The notebook change is a markdown cell, so its executed outputs are untouched. Docs gate: 0 errors, critical at baseline. Co-Authored-By: Claude Fable 5.1 --- CONTEXT.md | 6 ++++-- .../feature_identification.rst | 19 ++++++++++++------- use_cases/use_case1_gamma_secretase.ipynb | 7 +++++++ 3 files changed, 23 insertions(+), 9 deletions(-) diff --git a/CONTEXT.md b/CONTEXT.md index f667de62..89dc869f 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -29,9 +29,11 @@ _Avoid_: 0-based, half-open / exclusive-stop, `len()`-style stop (a stop equal t **part**: A named region of a protein over which a **split** operates and a scale is averaged; the `PART` field of a feature id (`PART-SPLIT-SCALE`). Parts are the columns of `df_parts`, produced by `SequenceFeature.get_df_parts`. -> **`tmd` means TARGET MIDDLE DOMAIN. It is already the general abstraction — do not "generalize" it.** +> **`tmd` was TRANSMEMBRANE DOMAIN and is now TARGET MIDDLE DOMAIN. It is already the general abstraction — do not "generalize" it.** > -> The part vocabulary is a **geometry**, not a biological claim: one **target middle domain** (`tmd`) flanked by two **juxta middle domains** (`jmd_n`, `jmd_c`), plus their composites (`jmd_n_tmd_n`, `tmd_c_jmd_c`, …). The letters are historical — the abstraction is "the span of interest and its two flanks", and it applies unchanged to a Pfam domain, a kinase domain, a cleavage-site window or a whole chain. A feature id reading `TMD-Segment(2,4)-ANDN920101` is therefore **not** claiming a transmembrane helix; it names the target span. +> The letters come from CPP's first application, to γ-secretase substrates, where the span of interest *is* the transmembrane domain and its flanks *are* the juxtamembrane domains. When the method was generalized the acronyms were **deliberately kept and re-read**: `tmd` as **target middle domain**, `jmd_n` / `jmd_c` as **juxta middle domains**. Same letters, same geometry, wider meaning. +> +> So the vocabulary is a **geometry**, not a biological claim: one target span flanked by two juxta spans, plus their composites (`jmd_n_tmd_n`, `tmd_c_jmd_c`, …), applying unchanged to a Pfam domain, a kinase domain, a cleavage-site window or a whole chain. A feature id reading `TMD-Segment(2,4)-ANDN920101` names the target span; it describes a transmembrane helix only when the data are membrane proteins. > > The published statement of this is the *Feature Identification* chapter of the docs > (`docs/source/index/usage_principles/feature_identification.rst`, anchor `part_vocabulary`), diff --git a/docs/source/index/usage_principles/feature_identification.rst b/docs/source/index/usage_principles/feature_identification.rst index 416e9624..f10f6001 100755 --- a/docs/source/index/usage_principles/feature_identification.rst +++ b/docs/source/index/usage_principles/feature_identification.rst @@ -30,15 +30,20 @@ The core idea of CPP is its feature concept: All possible parts are sub-parts or combinations of the **Target Middle Domain (TMD)**, **Juxta Middle Domain N-terminal (JMD-N)**, and **Juxta Middle Domain C-terminal (JMD-C)**. -.. admonition:: TMD means *Target Middle Domain*, not *transmembrane domain* +.. admonition:: TMD: from *Transmembrane Domain* to *Target Middle Domain* :class: important - The part vocabulary is a **geometry**, not a biological claim: one target span (``tmd``) - flanked by two juxta spans (``jmd_n``, ``jmd_c``), plus their composites. It is already - general and applies unchanged to a Pfam domain, a kinase domain, a cleavage-site window or - a whole chain. A feature id such as ``TMD-Segment(2,4)-ANDN920101`` therefore does **not** - assert a transmembrane helix; it names the target span. The letters are historical, as the - paragraph below explains, and there is nothing to rename per data set. + The letters come from CPP's first application, to substrates of γ-secretase, where the span + of interest **is** the transmembrane domain and the flanks **are** the juxtamembrane domains. + When the method was generalized beyond membrane proteins the acronyms were deliberately kept + and re-read: **TMD** as *Target Middle Domain*, **JMD** as *Juxta Middle Domain*. Same + letters, same geometry, wider meaning. + + So the vocabulary is a **geometry** - one target span flanked by two juxta spans - and it + already applies unchanged to a Pfam domain, a kinase domain, a cleavage-site window or a + whole chain. A feature id such as ``TMD-Segment(2,4)-ANDN920101`` names the target span; it + describes a transmembrane helix only when the data are membrane proteins, as in the + γ-secretase case below. There is nothing to rename per data set. .. figure:: /_artwork/schemes/scheme_CPP2.png diff --git a/use_cases/use_case1_gamma_secretase.ipynb b/use_cases/use_case1_gamma_secretase.ipynb index 4bf8c357..03634fc3 100644 --- a/use_cases/use_case1_gamma_secretase.ipynb +++ b/use_cases/use_case1_gamma_secretase.ipynb @@ -129,6 +129,13 @@ "- **`DOM_GSEC_PU`** — the **positive-unlabelled** set: the same 63 substrates + 631\n", " *unlabelled* proteins of **unknown** status (`label=2`, the \"others\").\n", "\n", + "**Where the names come from.** This analysis is the one CPP's part vocabulary is named after:\n", + "here `tmd` is literally the **transmembrane domain** and `jmd_n` / `jmd_c` are the\n", + "**juxtamembrane domains** flanking it. When CPP was generalized beyond membrane proteins the\n", + "acronyms were kept and re-read — **TMD** as *Target Middle Domain*, **JMD** as *Juxta Middle\n", + "Domain* — so the same three parts describe a Pfam domain or a cleavage-site window without\n", + "renaming anything. In this use case, read them in their original, literal sense.\n", + "\n", "**TMD annotation.** The `jmd_n` / `tmd` / `jmd_c` split depends on where each protein's\n", "**transmembrane domain (TMD)** sits — predicted from the sequence by a **TMD-annotation model**\n", "(which residues span the membrane, hence the boundaries defining the flanking JMD-N / JMD-C). The\n",