Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,19 @@ notes — with cross-references and examples — live in
rejected at construction with both counts in the message. Once the folds are consumed,
`df_folds_` reports per-fold sample counts, group counts and class balance. The package performs
the split; it runs no homology search or clustering itself (#478).
- `audit_leakage(df_seq=None, *, labels=None, groups=None, splits=None, X=None, names=None,
raise_on=None)`: inspects an evaluation setup for leakage risks and returns a plain `DataFrame`
of findings (`check`, `severity`, `detail`, `ids`), worst first, with the overall verdict in
`df_audit.attrs["status"]`. Without `splits` it audits the dataset, which is the check to run
before choosing a split; with `splits` it also compares the training against the test part
within each fold. It reports duplicate sequences, a row index, protein or group present on both
sides of a fold, a feature correlating with the label at `|r| >= 0.95`, an empty or strongly
uneven fold, and a single-class or skewed fold. Windowed input is handled correctly: `window`
takes precedence over the repeated parent `sequence`, so siblings are not mistaken for
duplicates. `raise_on` turns findings at or above a severity into a `ValueError`; the default
reports only and never raises. The checks are heuristic and a clean report is not a proof that
no leakage exists, and the severities are labels for a human reader, not a machine taxonomy
(#479).
- `Docs Build (gate)` CI workflow (`.github/scripts/check_docs_build.py`): builds the
documentation and fails on a docutils `ERROR`, so a broken reference cannot reach `master`
silently. It parses the build log rather than Sphinx's exit code, which is `0` even with
Expand Down
3 changes: 2 additions & 1 deletion aaanalysis/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
from .pu_learning import dPULearn, dPULearnPlot
from .explainable_ai import TreeModel
from .prediction import (AAPred, AAPredPlot, ReliabilityModel, ReliabilityModelPlot,
ModelEvaluator, ModelEvaluatorPlot, bind_groups)
ModelEvaluator, ModelEvaluatorPlot, bind_groups, audit_leakage)
from .protein_engineering import (DesignConstraints,
AAMut, AAMutPlot, SeqMut, SeqMutPlot, SeqOpt, SeqOptPlot)
from .plotting import (plot_get_clist, plot_get_cmap, plot_get_cdict,
Expand Down Expand Up @@ -74,6 +74,7 @@
"ModelEvaluator",
"ModelEvaluatorPlot",
"bind_groups",
"audit_leakage",
# "ShapModel" # SHAP
"plot_get_clist",
"plot_get_cmap",
Expand Down
27 changes: 27 additions & 0 deletions aaanalysis/_constants.py
Original file line number Diff line number Diff line change
Expand Up @@ -639,6 +639,33 @@ def _folder_path(super_folder, folder_name):
COLS_FOLDS_GROUPS = [COL_FOLD, COL_N_TRAIN, COL_N_TEST, COL_N_GROUPS_TRAIN, COL_N_GROUPS_TEST,
COL_POS_RATE_TRAIN, COL_POS_RATE_TEST]

# audit_leakage (heuristic leakage diagnostic): one row per finding, worst severity in .attrs
COL_CHECK = "check" # name of the heuristic that produced the finding
COL_SEVERITY = "severity" # human-readable label (not a machine taxonomy): low, medium, high
COL_DETAIL = "detail" # one-sentence, human-readable description of the finding
COL_IDS = "ids" # affected identifiers (list, capped; the true count is in detail)
COLS_AUDIT_LEAKAGE = [COL_CHECK, COL_SEVERITY, COL_DETAIL, COL_IDS]
# Severities, ordered from least to most severe. A label for a human reader, deliberately NOT a
# stable machine code: deciding whether a finding blocks a workflow is decision-layer policy.
STR_SEVERITY_LOW = "low"
STR_SEVERITY_MEDIUM = "medium"
STR_SEVERITY_HIGH = "high"
LIST_SEVERITIES = [STR_SEVERITY_LOW, STR_SEVERITY_MEDIUM, STR_SEVERITY_HIGH]
STR_STATUS_OK = "ok" # df.attrs["status"] when no finding was made
# Check names, i.e. the values of the 'check' column
STR_CHECK_DUPLICATE_SEQ = "duplicate_sequences" # identical sequence strings in df_seq
STR_CHECK_TRAIN_TEST_OVERLAP = "train_test_overlap" # a row index in both parts of a fold
STR_CHECK_DUPLICATE_SEQ_FOLDS = "duplicate_sequences_across_folds" # identical sequence split apart
STR_CHECK_ENTRY_FOLDS = "same_protein_across_folds" # windows of one protein split apart
STR_CHECK_GROUP_FOLDS = "group_overlap_across_folds" # a group id in both parts of a fold
STR_CHECK_TARGET_LEAK = "target_derived_feature" # a feature near-perfectly tracking the label
STR_CHECK_FOLD_SIZE = "fold_size_anomaly" # a test fold far from the average size
STR_CHECK_CLASS_BALANCE = "class_balance_anomaly" # a fold's class balance far from the whole
LIST_CHECKS_LEAKAGE = [STR_CHECK_DUPLICATE_SEQ, STR_CHECK_TRAIN_TEST_OVERLAP,
STR_CHECK_DUPLICATE_SEQ_FOLDS, STR_CHECK_ENTRY_FOLDS,
STR_CHECK_GROUP_FOLDS, STR_CHECK_TARGET_LEAK,
STR_CHECK_FOLD_SIZE, STR_CHECK_CLASS_BALANCE]

# Labels
LABEL_FEAT_VAL = "Feature value"
LABEL_HIST_COUNT = "Number of proteins"
Expand Down
9 changes: 7 additions & 2 deletions aaanalysis/prediction/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
Prediction: evaluate and deploy sequence-based prediction models.

Public objects: AAPred, AAPredPlot, ReliabilityModel, ReliabilityModelPlot, ModelEvaluator,
ModelEvaluatorPlot, bind_groups.
ModelEvaluatorPlot, bind_groups, audit_leakage.
Downstream of feature engineering (``CPP`` / ``CPPGrid`` produce ``df_feat`` and the feature
matrix ``X``): ``AAPred`` evaluates one or more scikit-learn models across metrics by
cross-validation and an optional held-out set (``eval``), and fits them for deployment, then
Expand All @@ -13,7 +13,10 @@
comparison with a signed delta and a Wilcoxon significance test (``eval``) — visualized by
``ModelEvaluatorPlot``; ``bind_groups`` binds group labels (protein accession, family, or an
externally computed homology cluster) to any scikit-learn splitter, so dependent samples stay
within one fold of ``eval`` / ``run`` instead of leaking across them. Complements ``explainable_ai.TreeModel`` (tree-ensemble feature
within one fold of ``eval`` / ``run`` instead of leaking across them, while ``audit_leakage``
inspects a dataset and its folds after the fact and reports the leakage risks it can see
(duplicate sequences, a protein or group split across a fold, a feature tracking the label, a
skewed fold) as a plain table of findings. Complements ``explainable_ai.TreeModel`` (tree-ensemble feature
importance) — this subpackage owns the general evaluate-and-deploy path.

See ``.claude/rules/code-conventions.md`` for conventions and ``CONTEXT.md`` for domain terms.
Expand All @@ -25,6 +28,7 @@
from ._model_evaluator import ModelEvaluator
from ._model_evaluator_plot import ModelEvaluatorPlot
from ._bind_groups import bind_groups
from ._audit_leakage import audit_leakage

__all__ = [
"AAPred",
Expand All @@ -34,4 +38,5 @@
"ModelEvaluator",
"ModelEvaluatorPlot",
"bind_groups",
"audit_leakage",
]
Loading
Loading