Skip to content

Metal: remove per-submit host allocations and round trips - #9

Merged
jcwal1516 merged 2 commits into
mainfrom
perf/metal-host-round-trips
Sep 27, 2026
Merged

jcwal1516 merged 2 commits into
mainfrom
perf/metal-host-round-trips

Conversation

@jcwal1516

Copy link
Copy Markdown
Member

Stacked on #8 (the overlap kernel entry points and encoder files overlap). Review that PR first; this PR's diff is only the host-side changes.

  • Overlap schedules. The CPU rebuilt them and uploaded them into a newly allocated MTLBuffer for every image and plane on every submission, and each work item baked in an absolute sample offset. Schedules are now built plane-relative, uploaded once per plane geometry into a bounded (64-entry) LRU cache on the runtime, and rebased by a per-dispatch base offset in the kernel. The host still rejects any rebased index beyond the u32 device ABI. A test checks that relative schedules plus the base equal the absolute schedules for soft and hard tiles in both passes.
  • Output descriptors. Small descriptor arrays are passed with setBytes, falling back to a buffer above 4 KiB, instead of one shared-buffer allocation per image.
  • Dense batches. Every image was encoded into its own command buffer on one queue, and Metal serialized them because they all write one tracked allocation. submit_dense_batch and submit_batch_into now group up to 16 images, within the existing 256 MB scratch budget, into concatenated-descriptor command buffers; batch outputs carry byte offsets for this. An oversized image still gets its own per-image command buffer, as before.
  • Host readback. Decoding 16- and 32-bit formats to host copied the shared output into a Vec<u8> and then into the typed vector; it now converts once from the mapped allocation. readback() and readback_batch_image() reuse one lazily created queue per session instead of creating a command queue per call.
  • Planning. JxrDecoder::decode reuses its routing plan for Metal and CUDA preparation instead of planning the request twice.
  • Batch width. With these host costs removed, 8 images per concurrent batch command gave higher 128-tile throughput than 2 in a sweep of 2/4/8 on the 16-core M4 Pro (≈330–340 vs ≈366–375 MP/s pipelined, ±5% run-to-run). The default is now 8.

Performance (M4 Pro, interleaved with main and #8 in the same session, better of two rounds). jxr-pathology-bench, 256×256 Boat tiles; "Metal-side" is submit + wait with CPU preparation excluded, and pipelined throughput includes CPU entropy decoding:

Batch Metric main #8 kernels this PR vs #8 vs main
8 pipelined throughput 160.8 MP/s 204.5 MP/s 247.7 MP/s +21% +54%
8 resident, Metal-side 1.9 ms 1.5 ms 0.8 ms -45% -56%
8 dense, Metal-side 2.9 ms 2.6 ms 1.0 ms -62% -66%
8 host, Metal-side 1.8 ms 1.6 ms 1.0 ms -38% -45%
32 pipelined throughput 225.7 MP/s 233.3 MP/s 304.1 MP/s +30% +35%
32 resident, Metal-side 4.5 ms 3.7 ms 3.0 ms -17% -33%
32 dense, Metal-side 8.7 ms 6.3 ms 4.5 ms -29% -48%
32 host, Metal-side 4.4 ms 3.5 ms 2.5 ms -30% -44%
128 pipelined throughput 241.6 MP/s 266.3 MP/s 349.1 MP/s +31% +44%
128 resident, Metal-side 15.7 ms 14.2 ms 8.5 ms -40% -46%
128 dense, Metal-side 31.5 ms 16.2 ms 9.3 ms -42% -70%
128 host, Metal-side 17.0 ms 15.2 ms 9.5 ms -38% -44%

Single-image Metal decodes (jxr-load-bench) change little beyond #8, because one image has no batching to amortize: cumulative against main, Seattle −11%, P19d −10%, Maui 128bpp −9%, Maui RGBE −7%.

Validation on an M4 Pro, on top of the kernels PR: cargo fmt --all -- --check; cargo clippy --workspace --all-targets --all-features -- -D warnings, and the same with --target x86_64-unknown-linux-gnu; cargo test --workspace --all-features (256 passed); cargo test -p jxr --no-default-features (24 passed); RUSTDOCFLAGS="-D warnings" cargo doc --workspace --all-features --no-deps; cargo test -p jxr-mpsgraph --test metal -- --ignored --test-threads=1 (7 passed, which exercises submit_batch_into on a caller queue); T.834/T.835 conformance Metal 517/517 and CPU 517/517; jxr-pathology-bench checksum validation of resident and dense batches. The owner-only metal-hardware workflow has not been run on this revision.

Every reconstruction kernel widened each add, subtract, and multiply to
64-bit, which Apple GPUs emulate, and returned through a branch after
each operation. Arithmetic now uses exact 32-bit overflow predicates
(sign-bit tests for add/sub, mulhi for quantizer products, range tests
for the times-three steps) accumulated into a sticky flag that each
kernel tests once before storing. The transforms become straight-line
code; per-phase status codes are unchanged, and results derived from an
overflowed intermediate are never stored.

The HP kernel ran one threadgroup per macroblock with thread 0 applying
HP prediction serially between two barriers, leaving half of each SIMD
group idle for 16-block luma macroblocks. It now runs one thread per
4x4 block in a flat grid; each thread accumulates its own prediction
chain in the normative order, so every partial-sum overflow check
matches the serial traversal. Block rows are stored as aligned int4
writes, and plane ABI construction now rejects sample planes that are
not four-sample aligned.

The output kernels converted color, including chroma upsampling, once
per output channel. Each pixel is now loaded and converted once, and
premultiplied stores scale alpha once per pixel. Chroma upsampling
computes the weighted average exactly in 32 bits by splitting each
operand into 8q + r, and unsigned premultiplication uses 32-bit
division (65535^2 + 32767 fits in u32).
Overlap schedules were rebuilt on the CPU and uploaded into a newly
allocated MTLBuffer for every image and plane on every submission, and
their work items baked in absolute sample offsets. Schedules are now
built plane-relative, uploaded once per plane geometry into a bounded
LRU cache on the runtime, and rebased in the kernel with a per-dispatch
base offset. The host still rejects any rebased index beyond the u32
device ABI.

Small output descriptor arrays are passed with setBytes instead of a
shared-buffer allocation per image.

Dense batches encoded every image into its own command buffer on one
queue; because all of them write one tracked allocation, Metal
serialized them. Dense and caller-destination batches now group up to
16 images, bounded by the batch scratch budget, into concatenated-
descriptor command buffers. Batch outputs carry byte offsets for this.

Host decodes of 16- and 32-bit formats copied the shared output into a
Vec<u8> and then into the typed vector; they now convert once from the
mapped allocation. Resident readback reuses one lazily created queue
per session instead of creating a command queue per call.
JxrDecoder::decode reuses its routing plan for Metal and CUDA
preparation instead of planning the request twice.

With these host costs removed, eight images per concurrent batch
command measured higher 128-tile pipeline throughput than two on the
16-core M4 Pro, so the batch width is now eight.
@jcwal1516
jcwal1516 changed the base branch from perf/metal-kernels to main September 27, 2026 00:31
@jcwal1516
jcwal1516 merged commit 846801d into main Sep 27, 2026
2 checks passed
@jcwal1516
jcwal1516 deleted the perf/metal-host-round-trips branch September 28, 2026 06:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant