Repository navigation
Add mafutils info and record block aggregates in the index header - #5
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
There was no cheap way to ask what a MAF contains.
statsreads the whole file andvalidateonly answers whether the index is trustworthy.mafutils infoprints size, compression, index freshness, block/scaffold counts and a species list. On a 42.70 GB file it takes 1.3 s (0.7 s with --sample-blocks 0), becausemafutils indexnow records the counts in the index header and info normally reads nothing but that one line.New header fields (blocks, scaffolds, ref_bases, aln_cols, seq_lines, max_seqs), appended after hash=. Written identically to both index files, since
validatecompares the two headers for exact equality. format stays at 2.Three tiers, never scanning the MAF for counts: header only; else stream the index (~52 MB/s) with a warning naming its size and recommending a rebuild; else report what needs no
indexand point atmafutils index.Species are sampled from the first 1,000 blocks and labelled possibly incomplete. They aren't in the
index, and collecting them during indexing measured 2.06× slower.statsremains the exhaustive source.Compatibility: index rows are byte-identical (verified across all 8,094,203 rows of a real index); v0.6.0 code reads a new-format index cleanly (fetch/stats/gc/validate all pass, validate VERIFIED); new code reads old indexes via the fallback. No rebuild required, it only makes info instant. Indexing cost measured at ~1.3% (4.71 s → 4.77 s on a 1.5 GB slice). A full 42.7 GB rebuild produced aggregates matching an independent
awkpass exactly. 148 tests passing.🤖 Generated with Claude Code