Public, resumable workflow for hyperrestrictive de novo crab transcriptome assembly, classical functional annotation, differential expression, and source-selectable enrichment.
The workflow turns paired-end RNA-seq libraries into a representative gene catalogue, predicted proteins, classical and protein-language-model functional annotations, expression matrices, differential-expression results, and source-specific enrichment tables. It was developed from three crab transcriptome projects while keeping sample count and species identity fully configurable.
- Hydra and Pydantic compose and validate portable scientific configuration.
uvlocks the Python control-plane environment.- Nextflow DSL2 executes preprocessing, Trinity, TransDecoder, Salmon, Trinotate, EggNOG-mapper, KofamScan, edgeR/DESeq2, and ORA.
- Grid'5000 profiles run only inside existing OAR allocations. Frontends are control-plane only.
The validated hyperrestrictive defaults are normalization max_cov=30, min_cov=3, paired reads together, and Trinity min_contig_length=800, min_kmer_cov=8, min_glue=8, min_iso_ratio=0.15, and group_pairs_distance=800.
The production assemblies for species01, species02, and species03 were run on
the Grid'5000 Nancy node grosminet-1, using its 72 CPU cores, large-memory
capacity, and node-local scratch. The portable profile g5k_grosminet
reproduces that resource strategy without making the workflow dependent on a
specific hostname. Real checkpoint-pinned ESM inference probes were validated
separately on the two NVIDIA H100 NVL GPUs of musa-2 at Sophia.
See scientific methods for the analysis stages and evidence boundaries, and validation evidence for the tests actually executed by this repository.
The workflow is species-agnostic and accepts any positive number of paired-end
libraries. One CSV row represents one biological library and must provide a
unique sample identifier, distinct existing left/right FASTQ files, and a
nonempty condition. Read files cannot be reused by another row. When edgeR or
DESeq2 is enabled, every condition named by a contrast must have at least two
replicates. See assets/samplesheet.example.csv.
FASTQ paths may be absolute or relative to the samplesheet directory. Sample
identifiers are validated before launch because they become durable output
names. See docs/CONFIGURATION.md and assets/contrasts.example.tsv for the
complete species-independent contract.
Use one fresh output directory per species and run. For example:
uv run crabtx preflight species=species01 paths.samplesheet=/data/species01.csv paths.contrasts=/data/species01_contrasts.tsv paths.outdir=/results/species01_run1
uv run crabtx preflight species=species02 paths.samplesheet=/data/species02.csv paths.contrasts=/data/species02_contrasts.tsv paths.outdir=/results/species02_run1
uv run crabtx preflight species=species03 paths.samplesheet=/data/species03.csv paths.contrasts=/data/species03_contrasts.tsv paths.outdir=/results/species03_run1Changing the number of rows never changes the hyperrestrictive scientific
parameters. It only changes the per-sample preprocessing and quantification
task universe.
Nextflow is pinned in .nextflow-version; version 26.04.6 supports the Java 25 runtime currently exposed by the tested Grid'5000 node.
Prerequisites are Linux or macOS, Java 17 or newer, Conda/Miniforge, uv, and
Nextflow. Install the two project-level launchers at their validated versions:
curl -LsSf https://astral.sh/uv/0.9.17/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
export NXF_VER="$(cat .nextflow-version)"
curl -s https://get.nextflow.io | bash
install -Dm755 nextflow "$HOME/.local/bin/nextflow"
uv --version
nextflow -version
conda --versionConda supplies the isolated bioinformatics environments declared in envs/;
the repository-level conf/condarc.yml restricts resolution to conda-forge
and Bioconda.
uv sync --frozen --group dev
uv run pytest
uv run crabtx config profile=test
uv run crabtx preflight profile=local paths.samplesheet=/absolute/samples.csv paths.outdir=/fresh/run
uv run crabtx run profile=local paths.samplesheet=/absolute/samples.csv paths.outdir=/fresh/runDifferential expression requires a tab-separated contrasts table with contrast, numerator, and denominator columns. Classical annotation databases follow the directory contract documented in docs/OUTPUTS.md and are checked before execution.
Use nextflow run main.nf -profile test -stub-run -params-file assets/test/params.json for an infrastructure-only smoke test. Production outputs are never overwritten; use Nextflow -resume only against a deliberately retained work directory.
CI runs Python tests, Ruff, nextflow lint ., and a complete CPU -stub-run.
Real-data acceptance is performed inside allocated Grid'5000 nodes. The four
configured ESM models have also passed checkpoint-pinned inference probes on
NVIDIA H100 GPUs. A complete GPU-profile -stub-run verifies that all four ESM
branches connect to the classical annotation, differential-expression, and
enrichment workflow. No CPU execution silently substitutes an ESM model.
Run uv run crabtx validate /path/to/run after completion to verify provenance, FASTA identifiers and counts, the gene-count matrix, requested classical annotation tables, differential-expression results, and the expected source-by-method enrichment table set.
See Grid'5000 execution, output contracts, and the design specification. CPU and H100 acceptance evidence is summarized in docs/VALIDATION.md.
| Document | Purpose |
|---|---|
| Scientific methods | Biological rationale, analysis stages, parameters, and execution provenance |
| Configuration | Samplesheets, contrasts, Hydra overrides, and species-specific runs |
| Outputs | Durable file contracts and reference acceptance counts |
| Validation | Automated checks, CPU probes, and H100 evidence |
| Grid'5000 | Safe OAR execution and node-local scratch usage |
| Contributing | Development, testing, and data-governance rules |
| Citation | Machine-readable software citation metadata |