Skip to content

Latest commit

 

History

49 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

faba

Feature statistics Accumulator for Base-pair-level Analysis

faba extracts per-cell genomic features directly from alignment (BAM) files: RNA modifications (DART-seq m6A, A-to-I editing), alternative polyadenylation (APA), gene counts, read depth, and SNP genotypes.

Its product is the matrices; nothing here fits a model. The embedding, trajectory and annotation commands that read them live in senna (legume-rs).

Installation

Prerequisites

These must be installed before running cargo install — Cargo only builds Rust crates and cannot pull in a C toolchain or system libraries on its own. If they are missing, the build fails midway in the vendored-htslib build script.

  • Rust (stable, edition 2021+) — install via rustup.

  • A C/C++ toolchain and the system libraries that rust-htslib builds against (it vendors and compiles htslib, which needs the compression dev headers, and runs bindgen, which needs libclang):

    # Debian / Ubuntu
    sudo apt-get install build-essential clang libclang-dev \
        zlib1g-dev libbz2-dev liblzma-dev pkg-config
    
    # macOS (Homebrew)
    brew install llvm xz bzip2 zlib

Install from crates.io

cargo install faba

Optional HDF5 (.h5 / .h5ad) support:

cargo install faba --features hdf5

Install from GitHub

cargo install --git https://github.com/causalpathlab/faba.git

Pin to a specific tag or branch if you need a reproducible build:

cargo install --git https://github.com/causalpathlab/faba.git --tag v0.14.0
cargo install --git https://github.com/causalpathlab/faba.git --branch main

Build from a local clone

git clone https://github.com/causalpathlab/faba.git
cd faba
cargo build --release
cargo install --path .

Optional features

Feature Enables Extra requirement
hdf5 HDF5 backend + .h5/.h5ad readers libhdf5 on the build host

There is no cuda / metal feature: nothing on the BAM-to-matrix path touches a GPU. senna (in legume-rs) carries those for the code that does.

Add them with --features:

cargo install --git https://github.com/causalpathlab/legume-rs.git faba --features hdf5

Verify

faba --help

Troubleshooting

Almost all install failures come from the hts-sys build step (it compiles a vendored copy of htslib and generates bindings with bindgen). The fixes:

  • thread 'main' panicked ... Unable to find libclang — bindgen can't locate libclang. Install libclang-dev (Debian/Ubuntu) or brew install llvm (macOS), then point to it:

    # Linux: the unversioned symlink ships with libclang-dev
    export LIBCLANG_PATH=$(llvm-config --libdir 2>/dev/null || echo /usr/lib/llvm-18/lib)
    
    # macOS (Homebrew llvm is keg-only)
    export LIBCLANG_PATH="$(brew --prefix llvm)/lib"
  • fatal error: zlib.h / bzlib.h / lzma.h: No such file or directory — the compression dev headers are missing (the runtime .so alone is not enough). Install zlib1g-dev libbz2-dev liblzma-dev (Debian/Ubuntu) or brew install xz bzip2 zlib (macOS).

  • error: linker 'cc' not found — no C toolchain. Install build-essential (Debian/Ubuntu) or Xcode Command Line Tools (xcode-select --install) on macOS.

  • failed to run custom build command for 'hdf5-sys' — you passed --features hdf5 without libhdf5 on the host. Install libhdf5-dev (Debian/Ubuntu) or brew install hdf5, or drop the feature.

If a build fails after fixing a prerequisite, re-run with a clean rebuild of the C bits: cargo install --git ... faba --force.

Usage

faba <COMMAND> [OPTIONS]
Command Purpose
Feature profiling — BAM → per-cell features
dartseq (dart, m6a) Call DART-seq m6A sites by a WT-vs-MUT control contrast on C-to-T conversions
atoi (a2i, editing) Detect and quantify A-to-I RNA editing sites
apa (polya) Quantify alternative polyadenylation sites per cell
count (genes) Count reads per gene and call cells (single-cell or bulk RNA-seq)
depth (rd) Compute read depth over genomic intervals
snp (genotype) Discover and genotype SNP variants from BAM pileup
all (pipeline) Run the full profiling pipeline: SNP → count → ATOI → m6A → APA
QC — choose and apply thresholds after profiling
qc-report Sweep every qc threshold and show kept sites / genes / cells, plus a -log10(p) histogram
qc Filter a faba output directory into a new fileset: cells, features and editing sites
Inspection & reference
pwm Build a position weight matrix around genomic sites
pileup (inspect) ASCII pileup, or a faceted Miami plot, for one gene
metagene (mg) Metagene histogram of site positions across gene features
docs Print the method write-ups compiled into this binary

Run faba <COMMAND> --help for the detailed options of each subcommand.

Embedding (gem), annotation (annotate-by-projection), trajectory (lineage, lineage-plot) and modality dynamics (dyn-assoc) are senna subcommands — they read the matrices above by prefix.

Methods

The method write-ups are compiled into the binary — include_str!, not files read at runtime — so they travel with a cargo installed faba or one copied to a cluster with no checkout beside it, and the build fails if one of them goes missing.

faba docs                # list what there is
faba docs profiling      # BAM -> per-cell features: m6A, A-to-I, APA, counts, SNPs

The annotation and lineage write-ups moved with their subcommands: senna docs.

The same files live in docs/, which also carries the design notes for work that is planned but not implemented (kept separate on purpose — reading a plan as though it described the code is how people end up debugging things that were never built).

Examples

# Gene counts and cell calling from a single-cell BAM
faba count sample.bam -g genes.gff -o out/

# A-to-I editing sites (every putative site is written; `faba qc` decides)
faba atoi sample.bam -g genes.gff -f genome.fa -o out/

# DART-seq m6A: signal (WT APOBEC1-YTH) vs catalytically-dead control (YTHmut).
# A control is REQUIRED — m6A can't be told apart from genomic C/T variation
# without it (--mut / --control / --background are accepted aliases).
faba dartseq wt.bam --control-bam ctrl.bam -g genes.gff -f genome.fa -o out/

# Everything in one pass (the m6A step runs only when --control-bam is given,
# otherwise it is skipped; the other steps need no control)
faba all sample.bam -g genes.gff -f genome.fa -o out/ --control-bam ctrl.bam

# The producers apply no p-value / effect-size / reproducibility cutoff.
# See what each threshold keeps, then cut into a NEW directory:
faba qc-report out/ -o out/qc
faba qc out/ -o out_qc/ --site-max-pv 0.05 --site-min-cells 10 --auto-cutoff

License

MIT — see the workspace license.

About

Feature statistics Accumulator for Base-pair-level Analysis (BAM → per-cell sparse matrices)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages