Skip to content

About

Reproducible hyperrestrictive de novo crab transcriptome assembly and functional annotation workflow using Hydra and Nextflow DSL2.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

crab-transcriptome

Public, resumable workflow for hyperrestrictive de novo crab transcriptome assembly, classical functional annotation, differential expression, and source-selectable enrichment.

The workflow turns paired-end RNA-seq libraries into a representative gene catalogue, predicted proteins, classical and protein-language-model functional annotations, expression matrices, differential-expression results, and source-specific enrichment tables. It was developed from three crab transcriptome projects while keeping sample count and species identity fully configurable.

Architecture

  • Hydra and Pydantic compose and validate portable scientific configuration.
  • uv locks the Python control-plane environment.
  • Nextflow DSL2 executes preprocessing, Trinity, TransDecoder, Salmon, Trinotate, EggNOG-mapper, KofamScan, edgeR/DESeq2, and ORA.
  • Grid'5000 profiles run only inside existing OAR allocations. Frontends are control-plane only.

The validated hyperrestrictive defaults are normalization max_cov=30, min_cov=3, paired reads together, and Trinity min_contig_length=800, min_kmer_cov=8, min_glue=8, min_iso_ratio=0.15, and group_pairs_distance=800.

Scientific provenance

The production assemblies for species01, species02, and species03 were run on the Grid'5000 Nancy node grosminet-1, using its 72 CPU cores, large-memory capacity, and node-local scratch. The portable profile g5k_grosminet reproduces that resource strategy without making the workflow dependent on a specific hostname. Real checkpoint-pinned ESM inference probes were validated separately on the two NVIDIA H100 NVL GPUs of musa-2 at Sophia.

See scientific methods for the analysis stages and evidence boundaries, and validation evidence for the tests actually executed by this repository.

Species and sample counts

The workflow is species-agnostic and accepts any positive number of paired-end libraries. One CSV row represents one biological library and must provide a unique sample identifier, distinct existing left/right FASTQ files, and a nonempty condition. Read files cannot be reused by another row. When edgeR or DESeq2 is enabled, every condition named by a contrast must have at least two replicates. See assets/samplesheet.example.csv.

FASTQ paths may be absolute or relative to the samplesheet directory. Sample identifiers are validated before launch because they become durable output names. See docs/CONFIGURATION.md and assets/contrasts.example.tsv for the complete species-independent contract.

Use one fresh output directory per species and run. For example:

uv run crabtx preflight species=species01 paths.samplesheet=/data/species01.csv paths.contrasts=/data/species01_contrasts.tsv paths.outdir=/results/species01_run1
uv run crabtx preflight species=species02 paths.samplesheet=/data/species02.csv paths.contrasts=/data/species02_contrasts.tsv paths.outdir=/results/species02_run1
uv run crabtx preflight species=species03 paths.samplesheet=/data/species03.csv paths.contrasts=/data/species03_contrasts.tsv paths.outdir=/results/species03_run1

Changing the number of rows never changes the hyperrestrictive scientific parameters. It only changes the per-sample preprocessing and quantification task universe. Nextflow is pinned in .nextflow-version; version 26.04.6 supports the Java 25 runtime currently exposed by the tested Grid'5000 node.

Quick start

Prerequisites are Linux or macOS, Java 17 or newer, Conda/Miniforge, uv, and Nextflow. Install the two project-level launchers at their validated versions:

curl -LsSf https://astral.sh/uv/0.9.17/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
export NXF_VER="$(cat .nextflow-version)"
curl -s https://get.nextflow.io | bash
install -Dm755 nextflow "$HOME/.local/bin/nextflow"
uv --version
nextflow -version
conda --version

Conda supplies the isolated bioinformatics environments declared in envs/; the repository-level conf/condarc.yml restricts resolution to conda-forge and Bioconda.

uv sync --frozen --group dev
uv run pytest
uv run crabtx config profile=test
uv run crabtx preflight profile=local paths.samplesheet=/absolute/samples.csv paths.outdir=/fresh/run
uv run crabtx run profile=local paths.samplesheet=/absolute/samples.csv paths.outdir=/fresh/run

Differential expression requires a tab-separated contrasts table with contrast, numerator, and denominator columns. Classical annotation databases follow the directory contract documented in docs/OUTPUTS.md and are checked before execution.

Use nextflow run main.nf -profile test -stub-run -params-file assets/test/params.json for an infrastructure-only smoke test. Production outputs are never overwritten; use Nextflow -resume only against a deliberately retained work directory.

Validation status

CI runs Python tests, Ruff, nextflow lint ., and a complete CPU -stub-run. Real-data acceptance is performed inside allocated Grid'5000 nodes. The four configured ESM models have also passed checkpoint-pinned inference probes on NVIDIA H100 GPUs. A complete GPU-profile -stub-run verifies that all four ESM branches connect to the classical annotation, differential-expression, and enrichment workflow. No CPU execution silently substitutes an ESM model.

Run uv run crabtx validate /path/to/run after completion to verify provenance, FASTA identifiers and counts, the gene-count matrix, requested classical annotation tables, differential-expression results, and the expected source-by-method enrichment table set.

See Grid'5000 execution, output contracts, and the design specification. CPU and H100 acceptance evidence is summarized in docs/VALIDATION.md.

Documentation

Document Purpose
Scientific methods Biological rationale, analysis stages, parameters, and execution provenance
Configuration Samplesheets, contrasts, Hydra overrides, and species-specific runs
Outputs Durable file contracts and reference acceptance counts
Validation Automated checks, CPU probes, and H100 evidence
Grid'5000 Safe OAR execution and node-local scratch usage
Contributing Development, testing, and data-governance rules
Citation Machine-readable software citation metadata

About

Reproducible hyperrestrictive de novo crab transcriptome assembly and functional annotation workflow using Hydra and Nextflow DSL2.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages