Skip to content

Read entropy bits a word at a time and decode VLCs by table - #6

Merged
jcwal1516 merged 2 commits into
mainfrom
perf/entropy-bit-reader
Sep 27, 2026
Merged

jcwal1516 merged 2 commits into
mainfrom
perf/entropy-bit-reader

Conversation

@jcwal1516

@jcwal1516 jcwal1516 commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

PacketBitReader::read_bits consumed one bit per loop iteration with a bounds-checked byte index. The three prefix-code decoders (entropy VLCs, CBPHP, and YUV DC/LP patterns) also read one bit at a time, linearly searching their code tables and recomputing the maximum code length on every call. Every backend runs this code: CPU, Metal, and CUDA all entropy-decode on the CPU.

Reads of up to 57 bits now take one big-endian 64-bit window load. End-of-buffer padding and error construction live in cold out-of-line paths, so the hot path inlines. Every normative prefix table compiles into an eight-bit lookup; construction is const, so an ambiguous or oversized table fails the build.

Error behavior is unchanged. Truncated packets report UnexpectedEnd at the bit where the serial decoder stopped, and unmatched full-length prefixes report InvalidVlc at the code start. Exhaustive tests compare each lookup decoder with the serial search for every 16-bit input, bit offset, and truncation length. They also compare the reader with a bit-serial reference for every width 0–64 at every offset.

Performance. Measured with jxr-load-bench (in-memory parse + entropy + reconstruction + packing/readback, 3 warmups, 20 iterations) on an M4 Pro, interleaved with main in the same session. Each value is the lower of two rounds' medians.

Image CPU main → branch Metal main → branch
P19d (800×534, 64bpp PRGBA) 36.8 → 30.9 ms (−16%) 31.3 → 25.2 ms (−19%)
Maui (1019×677, 128bpp fixed) 47.6 → 41.9 ms (−12%) 39.8 → 34.0 ms (−14%)
Maui (1019×677, 32bpp RGBE) 40.2 → 36.3 ms (−10%) 36.9 → 33.0 ms (−11%)
Seattle (800×531, RGB8) 27.1 → 24.9 ms (−8%) 23.4 → 21.2 ms (−10%)
VeryWideLevel255 (17152×128, RGB8) 53.6 → 55.3 ms (within noise) too noisy to quote

Gains track the number of entropy-coded bits: high-bit-depth images with many flexbits benefit most.

Validation on an M4 Pro: cargo fmt --all -- --check; cargo clippy --workspace --all-targets --all-features -- -D warnings; cargo test --workspace --all-features --release (259 passed); T.834/T.835 conformance CPU 517/517 and Metal 517/517.

Note: the first commit accidentally picked up workspace crate version bumps to 0.2.0 made in the working tree by something outside this change. The second commit restores every manifest and Cargo.lock to main, so the net diff is the six entropy files. Please squash-merge.

PacketBitReader::read_bits consumed one bit per loop iteration with a
bounds-checked byte index, and the three prefix-code decoders (entropy
VLCs, CBPHP, and YUV DC/LP patterns) read one bit at a time while
linearly searching their code tables and recomputing the maximum code
length on every call.

Reads of up to 57 bits now take one big-endian 64-bit window load, with
the end-of-buffer padding and error construction moved into cold
out-of-line paths so the hot path inlines. Every normative prefix table
is compiled into an eight-bit lookup; construction is const, so an
ambiguous or oversized table fails compilation.

Error behavior is unchanged: truncated packets report UnexpectedEnd at
the bit where the serial decoder stopped, and unmatched full-length
prefixes report InvalidVlc at the code start. Exhaustive tests compare
the lookup decoders with the serial search for every 16-bit input,
offset, and truncation, and the reader with a bit-serial reference for
every width and offset.
The previous commit unintentionally included workspace crate version
bumps to 0.2.0 and the matching lockfile entries. Version changes belong
to a release PR, so this restores every manifest and Cargo.lock to main.
@jcwal1516
jcwal1516 merged commit 46615ef into main Sep 27, 2026
2 checks passed
@jcwal1516
jcwal1516 deleted the perf/entropy-bit-reader branch September 28, 2026 06:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant