Metal: remove per-submit host allocations and round trips - #9
Merged
Merged
Conversation
Every reconstruction kernel widened each add, subtract, and multiply to 64-bit, which Apple GPUs emulate, and returned through a branch after each operation. Arithmetic now uses exact 32-bit overflow predicates (sign-bit tests for add/sub, mulhi for quantizer products, range tests for the times-three steps) accumulated into a sticky flag that each kernel tests once before storing. The transforms become straight-line code; per-phase status codes are unchanged, and results derived from an overflowed intermediate are never stored. The HP kernel ran one threadgroup per macroblock with thread 0 applying HP prediction serially between two barriers, leaving half of each SIMD group idle for 16-block luma macroblocks. It now runs one thread per 4x4 block in a flat grid; each thread accumulates its own prediction chain in the normative order, so every partial-sum overflow check matches the serial traversal. Block rows are stored as aligned int4 writes, and plane ABI construction now rejects sample planes that are not four-sample aligned. The output kernels converted color, including chroma upsampling, once per output channel. Each pixel is now loaded and converted once, and premultiplied stores scale alpha once per pixel. Chroma upsampling computes the weighted average exactly in 32 bits by splitting each operand into 8q + r, and unsigned premultiplication uses 32-bit division (65535^2 + 32767 fits in u32).
Overlap schedules were rebuilt on the CPU and uploaded into a newly allocated MTLBuffer for every image and plane on every submission, and their work items baked in absolute sample offsets. Schedules are now built plane-relative, uploaded once per plane geometry into a bounded LRU cache on the runtime, and rebased in the kernel with a per-dispatch base offset. The host still rejects any rebased index beyond the u32 device ABI. Small output descriptor arrays are passed with setBytes instead of a shared-buffer allocation per image. Dense batches encoded every image into its own command buffer on one queue; because all of them write one tracked allocation, Metal serialized them. Dense and caller-destination batches now group up to 16 images, bounded by the batch scratch budget, into concatenated- descriptor command buffers. Batch outputs carry byte offsets for this. Host decodes of 16- and 32-bit formats copied the shared output into a Vec<u8> and then into the typed vector; they now convert once from the mapped allocation. Resident readback reuses one lazily created queue per session instead of creating a command queue per call. JxrDecoder::decode reuses its routing plan for Metal and CUDA preparation instead of planning the request twice. With these host costs removed, eight images per concurrent batch command measured higher 128-tile pipeline throughput than two on the 16-core M4 Pro, so the batch width is now eight.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #8 (the overlap kernel entry points and encoder files overlap). Review that PR first; this PR's diff is only the host-side changes.
MTLBufferfor every image and plane on every submission, and each work item baked in an absolute sample offset. Schedules are now built plane-relative, uploaded once per plane geometry into a bounded (64-entry) LRU cache on the runtime, and rebased by a per-dispatchbaseoffset in the kernel. The host still rejects any rebased index beyond theu32device ABI. A test checks that relative schedules plus the base equal the absolute schedules for soft and hard tiles in both passes.setBytes, falling back to a buffer above 4 KiB, instead of one shared-buffer allocation per image.submit_dense_batchandsubmit_batch_intonow group up to 16 images, within the existing 256 MB scratch budget, into concatenated-descriptor command buffers; batch outputs carry byte offsets for this. An oversized image still gets its own per-image command buffer, as before.Vec<u8>and then into the typed vector; it now converts once from the mapped allocation.readback()andreadback_batch_image()reuse one lazily created queue per session instead of creating a command queue per call.JxrDecoder::decodereuses its routing plan for Metal and CUDA preparation instead of planning the request twice.Performance (M4 Pro, interleaved with
mainand #8 in the same session, better of two rounds).jxr-pathology-bench, 256×256 Boat tiles; "Metal-side" is submit + wait with CPU preparation excluded, and pipelined throughput includes CPU entropy decoding:Single-image Metal decodes (
jxr-load-bench) change little beyond #8, because one image has no batching to amortize: cumulative againstmain, Seattle −11%, P19d −10%, Maui 128bpp −9%, Maui RGBE −7%.Validation on an M4 Pro, on top of the kernels PR:
cargo fmt --all -- --check;cargo clippy --workspace --all-targets --all-features -- -D warnings, and the same with--target x86_64-unknown-linux-gnu;cargo test --workspace --all-features(256 passed);cargo test -p jxr --no-default-features(24 passed);RUSTDOCFLAGS="-D warnings" cargo doc --workspace --all-features --no-deps;cargo test -p jxr-mpsgraph --test metal -- --ignored --test-threads=1(7 passed, which exercisessubmit_batch_intoon a caller queue); T.834/T.835 conformance Metal 517/517 and CPU 517/517;jxr-pathology-benchchecksum validation of resident and dense batches. The owner-onlymetal-hardwareworkflow has not been run on this revision.