Skip to content

Faster reassign-cells scoring and pb tree folds; give back preload reservations - #25

Merged
YPARK merged 4 commits into
mainfrom
faster-reassign-preload-release
Oct 2, 2026
Merged

YPARK merged 4 commits into
mainfrom
faster-reassign-preload-release

Conversation

@YPARK

@YPARK YPARK commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Two speedups to the pseudobulk steps, with identical results, plus a fix that gives back preload memory.

Faster DC-Poisson scoring (alg/dc_poisson.rs)

DcPoissonStats::gene_sum / log_gene are now stored gene-major (g*K + k, with a new idx(k, g) helper). They used to be group-major. compute_log_probs_restricted scores all allowed groups in one pass over the entity's nonzero genes, reading K values next to each other for each gene. It skips the K×M table stride it used to take for every group.

Each group's sum still adds the same terms in the same order (f32 cache → f64, in gene order). The leave-one-out score for the current group is computed exactly as before. So scores are bit-for-bit identical. A new test checks this with to_bits() against the per-group reference, on full and partial candidate sets, after moves. Partial candidate lists are scored by their position in the list, using a scratch buffer for each thread. If a list repeats a group, each copy gets the same correct score, just as in the old per-group scoring. A test covers this.

Public API: the fields are still pub, but their memory layout has changed. Nothing outside this file reads them: I grepped data-beans, senna-rs and pinto-rs. pinto only uses the refine_* / proposer API, which has not changed.

pb tree folds (collapse_data/pb_tree.rs)

NodeProfiles::new, residual_variance and side_sums fold into dense accumulators that cover every gene. Rayon used to split small leaves into many tasks of a few cells each, and each task zeroed and merged 36k-long vectors. Now each task takes at least 256 cells (with_min_len). The math is unchanged. Rayon's fold grouping was already decided at run time, so this does not make the output any less deterministic.

Preload reservation (sparse_io/helpers.rs, zarr + hdf5 backends)

  • New reserve_preload(nnz, what) -> Option<PreloadReservation>. The returned RAII guard is stored next to the preloaded arrays. It gives back its share on clean_preloaded_* and when the backend is dropped. A clone reserves its share again.
  • preload_columns / preload_rows now do nothing if the data is already preloaded.
  • preload_within_budget stays for compatibility (it keeps its reservation, as before). preload_reserved_bytes() reports the bytes reserved now.
  • Tests: guard release and refusal (unit), and a backend test that runs preload → repeat → clean → drop on zarr and hdf5.

Measured

Data: 10k_BMMNC (8,195 cells × 36,591 genes, 128 groups, 16 threads). Each figure is the wall time between log lines, from 3 runs per binary of senna topic --pb-tree refined:

step before after
reassign cells (includes its prep) 7.2 / 7.2 / 9.2 s 3.4 / 3.5 / 4.2 s
pb tree 8.1 / 8.1 / 7.9 s 1.4 / 2.1 / 1.6 s

Identical results: reassign cells: 14438 moves in every run. pb_tree.json, cell_to_pb.parquet, pb_gene.parquet, coarsening.json and pb_reference.json have the same md5 before and after.

Checks: cargo test --features sim passes. Clippy reports no lints in the changed files. The -D warnings errors it shows are new rust 1.98 lints in files this PR does not touch, and are already present on main.

@YPARK
YPARK merged commit 0d0caaa into main Oct 2, 2026
5 checks passed
@YPARK
YPARK deleted the faster-reassign-preload-release branch October 2, 2026 21:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant