Faster reassign-cells scoring and pb tree folds; give back preload reservations - #25
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two speedups to the pseudobulk steps, with identical results, plus a fix that gives back preload memory.
Faster DC-Poisson scoring (
alg/dc_poisson.rs)DcPoissonStats::gene_sum/log_geneare now stored gene-major (g*K + k, with a newidx(k, g)helper). They used to be group-major.compute_log_probs_restrictedscores all allowed groups in one pass over the entity's nonzero genes, reading K values next to each other for each gene. It skips the K×M table stride it used to take for every group.Each group's sum still adds the same terms in the same order (f32 cache → f64, in gene order). The leave-one-out score for the current group is computed exactly as before. So scores are bit-for-bit identical. A new test checks this with
to_bits()against the per-group reference, on full and partial candidate sets, after moves. Partial candidate lists are scored by their position in the list, using a scratch buffer for each thread. If a list repeats a group, each copy gets the same correct score, just as in the old per-group scoring. A test covers this.Public API: the fields are still
pub, but their memory layout has changed. Nothing outside this file reads them: I grepped data-beans, senna-rs and pinto-rs. pinto only uses therefine_*/ proposer API, which has not changed.pb tree folds (
collapse_data/pb_tree.rs)NodeProfiles::new,residual_varianceandside_sumsfold into dense accumulators that cover every gene. Rayon used to split small leaves into many tasks of a few cells each, and each task zeroed and merged 36k-long vectors. Now each task takes at least 256 cells (with_min_len). The math is unchanged. Rayon's fold grouping was already decided at run time, so this does not make the output any less deterministic.Preload reservation (
sparse_io/helpers.rs, zarr + hdf5 backends)reserve_preload(nnz, what) -> Option<PreloadReservation>. The returned RAII guard is stored next to the preloaded arrays. It gives back its share onclean_preloaded_*and when the backend is dropped. A clone reserves its share again.preload_columns/preload_rowsnow do nothing if the data is already preloaded.preload_within_budgetstays for compatibility (it keeps its reservation, as before).preload_reserved_bytes()reports the bytes reserved now.Measured
Data: 10k_BMMNC (8,195 cells × 36,591 genes, 128 groups, 16 threads). Each figure is the wall time between log lines, from 3 runs per binary of
senna topic --pb-tree refined:Identical results:
reassign cells: 14438 movesin every run.pb_tree.json,cell_to_pb.parquet,pb_gene.parquet,coarsening.jsonandpb_reference.jsonhave the same md5 before and after.Checks:
cargo test --features simpasses. Clippy reports no lints in the changed files. The-D warningserrors it shows are new rust 1.98 lints in files this PR does not touch, and are already present on main.