Skip to content

Resolve output scaling and pixel layout once per image - #7

Merged
jcwal1516 merged 1 commit into
mainfrom
perf/cpu-output-packing
Sep 27, 2026
Merged

jcwal1516 merged 1 commit into
mainfrom
perf/cpu-output-packing

Conversation

@jcwal1516

Copy link
Copy Markdown
Member

The generic CPU packers interpreted the output format per sample:

  • every sample re-derived its bias, rounding, and postscaling from the request;
  • every pixel re-converted the crop origin, re-matched the channel layout, and copied its primaries with a variable-length memmove;
  • the packed RGB555/RGB565/RGB101010/RGBE path repeated the same per-sample scale resolution.

scale_integer_component is now "resolve a ChannelScale, then apply it". The ordered and packed-color packers resolve each channel's scale once per image through that same implementation, so the fast path and the reference cannot diverge. Resolution failures are stored and reported by the first sample that uses the channel, which preserves the previous error behavior. Crop conversion and layout decisions are hoisted out of the pixel loops, and primaries are copied with fixed-size moves.

This covers every non-Luma-U8 CPU output, including the common RGB8/RGBA8 targets. Luma U8 already had its own SIMD path.

Performance. Measured with jxr-load-bench (in-memory parse + entropy + reconstruction + packing/readback, 3 warmups, 20 iterations) on an M4 Pro, interleaved with main in the same session. Each value is the lower of two rounds' medians. CPU decode:

Image main → branch
VeryWideLevel255 (17152×128, RGB8) 53.6 → 41.8 ms (−22%)
Seattle (800×531, RGB8) 27.1 → 24.6 ms (−9%)
Maui (1019×677, 128bpp fixed) 47.6 → 44.0 ms (−8%)
P19d (800×534, 64bpp PRGBA) 36.8 → 34.1 ms (−7%)
Maui (1019×677, 32bpp RGBE) 40.2 → 40.6 ms (within noise)

Metal decodes don't use this path and are unchanged.

Validation on an M4 Pro: cargo fmt --all -- --check; cargo clippy --workspace --all-targets --all-features -- -D warnings; cargo test --workspace --all-features (255 passed); T.834/T.835 conformance CPU 517/517, Annex-A rewrap 517/517, and Metal 517/517. An earlier revision regressed RGBE by about 3.5%, because the packed-color path still resolved scales per sample. That's fixed here: RGBE, F16 (Maui-64bppRGBHalf), and F32 (P19d-128bppPRGBAFloat) CPU decodes match main within noise in same-session A/B runs.

The generic CPU packers interpreted the output format per sample: every
sample re-derived its bias, rounding, and postscaling from the request,
every pixel re-converted the crop origin and re-matched the channel
layout, and every pixel copied its primaries with a variable-length
memmove. The packed RGB555/RGB565/RGB101010/RGBE path repeated the same
per-sample scale resolution.

scale_integer_component is now "resolve a ChannelScale, then apply it",
and the ordered and packed-color packers resolve each channel's scale
once per image through that same implementation, so the fast path and
the reference cannot diverge. Resolution failures are kept and reported
by the first sample that uses the channel, preserving the previous
error behavior. Crop conversion and layout decisions are hoisted out of
the pixel loops, and primaries are copied with fixed-size moves.
@jcwal1516
jcwal1516 merged commit d4fdc58 into main Sep 27, 2026
2 checks passed
@jcwal1516
jcwal1516 deleted the perf/cpu-output-packing branch September 28, 2026 06:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant