diff --git a/CONTEXT.md b/CONTEXT.md index f667de62..89dc869f 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -29,9 +29,11 @@ _Avoid_: 0-based, half-open / exclusive-stop, `len()`-style stop (a stop equal t **part**: A named region of a protein over which a **split** operates and a scale is averaged; the `PART` field of a feature id (`PART-SPLIT-SCALE`). Parts are the columns of `df_parts`, produced by `SequenceFeature.get_df_parts`. -> **`tmd` means TARGET MIDDLE DOMAIN. It is already the general abstraction — do not "generalize" it.** +> **`tmd` was TRANSMEMBRANE DOMAIN and is now TARGET MIDDLE DOMAIN. It is already the general abstraction — do not "generalize" it.** > -> The part vocabulary is a **geometry**, not a biological claim: one **target middle domain** (`tmd`) flanked by two **juxta middle domains** (`jmd_n`, `jmd_c`), plus their composites (`jmd_n_tmd_n`, `tmd_c_jmd_c`, …). The letters are historical — the abstraction is "the span of interest and its two flanks", and it applies unchanged to a Pfam domain, a kinase domain, a cleavage-site window or a whole chain. A feature id reading `TMD-Segment(2,4)-ANDN920101` is therefore **not** claiming a transmembrane helix; it names the target span. +> The letters come from CPP's first application, to γ-secretase substrates, where the span of interest *is* the transmembrane domain and its flanks *are* the juxtamembrane domains. When the method was generalized the acronyms were **deliberately kept and re-read**: `tmd` as **target middle domain**, `jmd_n` / `jmd_c` as **juxta middle domains**. Same letters, same geometry, wider meaning. +> +> So the vocabulary is a **geometry**, not a biological claim: one target span flanked by two juxta spans, plus their composites (`jmd_n_tmd_n`, `tmd_c_jmd_c`, …), applying unchanged to a Pfam domain, a kinase domain, a cleavage-site window or a whole chain. A feature id reading `TMD-Segment(2,4)-ANDN920101` names the target span; it describes a transmembrane helix only when the data are membrane proteins. > > The published statement of this is the *Feature Identification* chapter of the docs > (`docs/source/index/usage_principles/feature_identification.rst`, anchor `part_vocabulary`), diff --git a/docs/source/index/usage_principles/feature_identification.rst b/docs/source/index/usage_principles/feature_identification.rst index 416e9624..f10f6001 100755 --- a/docs/source/index/usage_principles/feature_identification.rst +++ b/docs/source/index/usage_principles/feature_identification.rst @@ -30,15 +30,20 @@ The core idea of CPP is its feature concept: All possible parts are sub-parts or combinations of the **Target Middle Domain (TMD)**, **Juxta Middle Domain N-terminal (JMD-N)**, and **Juxta Middle Domain C-terminal (JMD-C)**. -.. admonition:: TMD means *Target Middle Domain*, not *transmembrane domain* +.. admonition:: TMD: from *Transmembrane Domain* to *Target Middle Domain* :class: important - The part vocabulary is a **geometry**, not a biological claim: one target span (``tmd``) - flanked by two juxta spans (``jmd_n``, ``jmd_c``), plus their composites. It is already - general and applies unchanged to a Pfam domain, a kinase domain, a cleavage-site window or - a whole chain. A feature id such as ``TMD-Segment(2,4)-ANDN920101`` therefore does **not** - assert a transmembrane helix; it names the target span. The letters are historical, as the - paragraph below explains, and there is nothing to rename per data set. + The letters come from CPP's first application, to substrates of γ-secretase, where the span + of interest **is** the transmembrane domain and the flanks **are** the juxtamembrane domains. + When the method was generalized beyond membrane proteins the acronyms were deliberately kept + and re-read: **TMD** as *Target Middle Domain*, **JMD** as *Juxta Middle Domain*. Same + letters, same geometry, wider meaning. + + So the vocabulary is a **geometry** - one target span flanked by two juxta spans - and it + already applies unchanged to a Pfam domain, a kinase domain, a cleavage-site window or a + whole chain. A feature id such as ``TMD-Segment(2,4)-ANDN920101`` names the target span; it + describes a transmembrane helix only when the data are membrane proteins, as in the + γ-secretase case below. There is nothing to rename per data set. .. figure:: /_artwork/schemes/scheme_CPP2.png diff --git a/use_cases/use_case1_gamma_secretase.ipynb b/use_cases/use_case1_gamma_secretase.ipynb index 4bf8c357..03634fc3 100644 --- a/use_cases/use_case1_gamma_secretase.ipynb +++ b/use_cases/use_case1_gamma_secretase.ipynb @@ -129,6 +129,13 @@ "- **`DOM_GSEC_PU`** — the **positive-unlabelled** set: the same 63 substrates + 631\n", " *unlabelled* proteins of **unknown** status (`label=2`, the \"others\").\n", "\n", + "**Where the names come from.** This analysis is the one CPP's part vocabulary is named after:\n", + "here `tmd` is literally the **transmembrane domain** and `jmd_n` / `jmd_c` are the\n", + "**juxtamembrane domains** flanking it. When CPP was generalized beyond membrane proteins the\n", + "acronyms were kept and re-read — **TMD** as *Target Middle Domain*, **JMD** as *Juxta Middle\n", + "Domain* — so the same three parts describe a Pfam domain or a cleavage-site window without\n", + "renaming anything. In this use case, read them in their original, literal sense.\n", + "\n", "**TMD annotation.** The `jmd_n` / `tmd` / `jmd_c` split depends on where each protein's\n", "**transmembrane domain (TMD)** sits — predicted from the sequence by a **TMD-annotation model**\n", "(which residues span the membrane, hence the boundaries defining the flanking JMD-N / JMD-C). The\n",