Skip to content

Release JXR 0.2.0 with CPU and Metal improvements - #10

Merged
jcwal1516 merged 13 commits into
mainfrom
release/0.2.0
Sep 27, 2026
Merged

jcwal1516 merged 13 commits into
mainfrom
release/0.2.0

Conversation

@jcwal1516

@jcwal1516 jcwal1516 commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

Release JXR 0.2.0 with the J2K 0.11 public-type upgrade and the integrated CPU entropy decoding, SIMD output packing, Metal kernel, and host-submission improvements from #5–#9. All seven affected public crates move to 0.2.0; jxr-math remains at its published 0.1.0 version.

Pin j2k-core, j2k-metal-support, and j2k-mpsgraph-support to the published 0.11.2 release used by wsi-rs. Both the workspace and fuzz lockfiles resolve J2K from crates.io, without sibling-checkout overrides.

Package contents have been reviewed and the jxr-core publication dry run passed. Hosted CI and CPU/Metal hardware validation run on this PR; CUDA hardware validation runs from main after merge. Publish in the dependency order documented in docs/releasing.md only after the release checks pass.

PacketBitReader::read_bits consumed one bit per loop iteration with a
bounds-checked byte index, and the three prefix-code decoders (entropy
VLCs, CBPHP, and YUV DC/LP patterns) read one bit at a time while
linearly searching their code tables and recomputing the maximum code
length on every call.

Reads of up to 57 bits now take one big-endian 64-bit window load, with
the end-of-buffer padding and error construction moved into cold
out-of-line paths so the hot path inlines. Every normative prefix table
is compiled into an eight-bit lookup; construction is const, so an
ambiguous or oversized table fails compilation.

Error behavior is unchanged: truncated packets report UnexpectedEnd at
the bit where the serial decoder stopped, and unmatched full-length
prefixes report InvalidVlc at the code start. Exhaustive tests compare
the lookup decoders with the serial search for every 16-bit input,
offset, and truncation, and the reader with a bit-serial reference for
every width and offset.
The vectorized U8 packer and HP coefficient scaler each validated their
input in a separate full pass before the vector loop: the packer
re-read every sample to check the output-bias addition, and the scaler
widened every coefficient to i64 to check the product.

Both checks now run inside the existing vector loop by tracking the
observed sample range with fearless_simd 1.0's lane max/min reductions,
then comparing it with the representable range once. Dequantization
compares against [i32::MIN / step, i32::MAX / step], which is exact for
every nonzero step. The packer bounds only its final row, since row
starts increase monotonically. Out-of-range input still returns the
same error; any wrapped values already written are discarded with the
failed call. Scalar paths saturate or wrap so debug builds cannot
overflow before the check.

Microbenchmarks on an M4 Pro (NEON): 256-coefficient dequantization
96.4 -> 21.0 ns per call, and 1-megapixel luma U8 packing 361 -> 133 us.
The generic CPU packers interpreted the output format per sample: every
sample re-derived its bias, rounding, and postscaling from the request,
every pixel re-converted the crop origin and re-matched the channel
layout, and every pixel copied its primaries with a variable-length
memmove. The packed RGB555/RGB565/RGB101010/RGBE path repeated the same
per-sample scale resolution.

scale_integer_component is now "resolve a ChannelScale, then apply it",
and the ordered and packed-color packers resolve each channel's scale
once per image through that same implementation, so the fast path and
the reference cannot diverge. Resolution failures are kept and reported
by the first sample that uses the channel, preserving the previous
error behavior. Crop conversion and layout decisions are hoisted out of
the pixel loops, and primaries are copied with fixed-size moves.
Every reconstruction kernel widened each add, subtract, and multiply to
64-bit, which Apple GPUs emulate, and returned through a branch after
each operation. Arithmetic now uses exact 32-bit overflow predicates
(sign-bit tests for add/sub, mulhi for quantizer products, range tests
for the times-three steps) accumulated into a sticky flag that each
kernel tests once before storing. The transforms become straight-line
code; per-phase status codes are unchanged, and results derived from an
overflowed intermediate are never stored.

The HP kernel ran one threadgroup per macroblock with thread 0 applying
HP prediction serially between two barriers, leaving half of each SIMD
group idle for 16-block luma macroblocks. It now runs one thread per
4x4 block in a flat grid; each thread accumulates its own prediction
chain in the normative order, so every partial-sum overflow check
matches the serial traversal. Block rows are stored as aligned int4
writes, and plane ABI construction now rejects sample planes that are
not four-sample aligned.

The output kernels converted color, including chroma upsampling, once
per output channel. Each pixel is now loaded and converted once, and
premultiplied stores scale alpha once per pixel. Chroma upsampling
computes the weighted average exactly in 32 bits by splitting each
operand into 8q + r, and unsigned premultiplication uses 32-bit
division (65535^2 + 32767 fits in u32).
Overlap schedules were rebuilt on the CPU and uploaded into a newly
allocated MTLBuffer for every image and plane on every submission, and
their work items baked in absolute sample offsets. Schedules are now
built plane-relative, uploaded once per plane geometry into a bounded
LRU cache on the runtime, and rebased in the kernel with a per-dispatch
base offset. The host still rejects any rebased index beyond the u32
device ABI.

Small output descriptor arrays are passed with setBytes instead of a
shared-buffer allocation per image.

Dense batches encoded every image into its own command buffer on one
queue; because all of them write one tracked allocation, Metal
serialized them. Dense and caller-destination batches now group up to
16 images, bounded by the batch scratch budget, into concatenated-
descriptor command buffers. Batch outputs carry byte offsets for this.

Host decodes of 16- and 32-bit formats copied the shared output into a
Vec<u8> and then into the typed vector; they now convert once from the
mapped allocation. Resident readback reuses one lazily created queue
per session instead of creating a command queue per call.
JxrDecoder::decode reuses its routing plan for Metal and CUDA
preparation instead of planning the request twice.

With these host costs removed, eight images per concurrent batch
command measured higher 128-tile pipeline throughput than two on the
16-core M4 Pro, so the batch width is now eight.
The previous commit unintentionally included workspace crate version
bumps to 0.2.0 and the matching lockfile entries. Version changes belong
to a release PR, so this restores every manifest and Cargo.lock to main.
jxr-core re-exports j2k_core::{BackendKind, BackendRequest, Rect}, so
the J2K 0.10 -> 0.11 upgrade changes public types. Every published crate
that exposes them moves to 0.2.0; jxr-math stays at 0.1.0.

docs/releasing.md: restore the published 0.1.2 record (jxr-mpsgraph 0.1.1
shipped with j2k-mpsgraph-support =0.11.0) and add the 0.2.0 section.
@jcwal1516
jcwal1516 changed the base branch from chore/deps-j2k-0.11.1 to main September 27, 2026 00:28
@jcwal1516 jcwal1516 changed the title build: release jxr 0.2.0 crate versions Release JXR 0.2.0 with CPU and Metal improvements Sep 27, 2026
@jcwal1516
jcwal1516 merged commit 06061e0 into main Sep 27, 2026
3 checks passed
@jcwal1516
jcwal1516 deleted the release/0.2.0 branch September 28, 2026 06:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant