Conversation
(cherry picked from commit a2a09fd)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Benchmark: Gemma 4 26B A4B runner, control against this stackResult: no measurable change. This is the expected result. The runner uses Setup
Prefill rate is prompt tokens divided by TTFT (client clock). Decode rate is from stream chunk times. Peak memory is the server process physical footprint (n = 2 server starts per arm). Long prompt (16,383 prompt tokens, 1,024 output tokens)
Read: the differences are smaller than the spread of the control arm (prefill 3943 to 4511 tok/s; peak memory 21.43 and 22.49 GiB across its two starts). In round 2 the arms agree to 0.4% on prefill and to 0.001 GiB on memory. Short prompt (2,059 prompt tokens, 642 output tokens)
Read: no change. Medians are given because the first request of round 1 was slow in the two arms (TTFT 1.19 s and 1.22 s). OutputThe output text hash is the same in all 12 runs of each prompt, in the two arms. Limit of this benchmarkThis benchmark shows that the pin move does not change a |
Benchmark: Qwen 3.5 35B A3B runner, control against this stackResult: no measurable change. This is the expected result. The runner uses Setup
The builds and the prompts are the same as in the Gemma benchmark. Prefill rate is prompt tokens divided by TTFT (client clock). TTFT is the time to the first streamed delta, reasoning text included. Decode rate is from stream chunk times. Values are medians, because a small number of runs were slow in each arm. Long prompt (14,666 prompt tokens, 1,024 output tokens)
Read: the difference comes from round 1 only, where the treatment runs were slower (prefill 3501 to 3694 tok/s). In round 2 the arms agree to 0.3% on prefill (3915 and 3926 tok/s, mean) and to 0.6% on decode. The sign is opposite to the Gemma long-prompt result (+2.31%). The two results are consistent with run-to-run noise. Short prompt (1,886 prompt tokens, 1,024 output tokens)
Read: no change in prefill, TTFT, or memory. The decode difference (0.8%) is of the same size as the difference between the two control rounds (130.3 and 129.1 tok/s, mean). Process peak memoryThe lifetime maximum of the server footprint is 30.94 GiB in the control arm and 30.95 GiB in the treatment arm (+0.02%). This maximum occurs before the measured runs (model load and warmup). OutputThe output text hash is the same in all 12 runs of each prompt, in the two arms. Limit of this benchmarkThis benchmark shows that the pin move does not change a |
(cherry picked from commit 9008944)
Benchmark with the
|
| Item | Value |
|---|---|
| Machine | Apple M5 Max, 128 GiB, macOS 26.6.2 |
| Models | mlx-community/gemma-4-26B-A4B-it-qat-4bit; EigenLabs/Qwen3.5-35B-A3B-MLX-VL-4bit-g64 at 8964653 |
| Control | d-inference 4d679360 (mlx 3fa8f25e) |
| Treatment | d-inference ff1eb060 (mlx aa7af3eb, mlx-swift dee30fb): ml-explore#4567 and ml-explore#4572 |
| Server | darkbloom start --local, MTP off, 1 concurrent request, contiguous KV |
| Sampling | greedy, seed 20260928, max 1,024 output tokens |
| Design | 2 interleaved rounds (control, treatment), 1 server start per arm per round, 3 measured runs per prompt. n = 6 per cell. |
| Conditions | GPU lock held, resident service unloaded, one window for the two models |
Prefill rate is prompt tokens divided by TTFT (client clock). TTFT is the time to the first streamed delta, reasoning text included. Decode rate is from stream chunk times. Values are medians. Peak memory is the server process physical footprint in the measured runs (n = 2 server starts per arm).
Qwen 3.5 35B A3B
| Prompt | Metric | Control | Treatment | Δ treatment − control |
|---|---|---|---|---|
| Long (14,666 tokens) | Prefill (tok/s) | 3891 | 4143 | +6.49% |
| Long | Decode (tok/s) | 117.4 | 119.1 | +1.42% |
| Long | TTFT (s) | 3.770 | 3.540 | −6.09% |
| Long | Wall time (s) | 12.50 | 12.13 | −2.95% |
| Long | Peak memory (GiB) | 27.63 | 27.11 | −1.88% |
| Short (1,886 tokens) | Prefill (tok/s) | 4570 | 4726 | +3.41% |
| Short | Decode (tok/s) | 129.5 | 131.4 | +1.50% |
| Short | TTFT (s) | 0.4127 | 0.3990 | −3.32% |
| Short | Wall time (s) | 8.312 | 8.180 | −1.59% |
| Short | Peak memory (GiB) | 21.96 | 21.96 | −0.00% |
Read: all speed metrics improve, and the gains repeat in the two rounds. On the long prompt, each treatment run is faster than the control run in the same position (run 1: 3817 and 3827 against 3549 and 3594 tok/s; runs 2 and 3: 4129 to 4168 against 3870 to 3931 tok/s). On the short prompt, the decode ranges do not overlap (treatment 130.3 to 132.0, control 128.9 to 130.1 tok/s). The long-prompt memory difference comes from one control start (28.14 GiB); the other three starts are 27.10 to 27.12 GiB.
Gemma 4 26B A4B
| Prompt | Metric | Control | Treatment | Δ treatment − control |
|---|---|---|---|---|
| Long (16,383 tokens) | Prefill (tok/s) | 4299 | 4393 | +2.20% |
| Long | Decode (tok/s) | 92.60 | 93.82 | +1.32% |
| Long | TTFT (s) | 3.813 | 3.729 | −2.19% |
| Long | Wall time (s) | 14.87 | 14.64 | −1.57% |
| Long | Peak memory (GiB) | 22.46 | 22.46 | +0.04% |
| Short (2,059 tokens) | Prefill (tok/s) | 4921 | 5045 | +2.52% |
| Short | Decode (tok/s) | 108.1 | 109.4 | +1.26% |
| Short | TTFT (s) | 0.4184 | 0.4082 | −2.44% |
| Short | Wall time (s) | 6.345 | 6.274 | −1.12% |
| Short | Peak memory (GiB) | 17.05 | 17.06 | +0.05% |
Read: the decode gain is measurable. The decode ranges do not overlap on the two prompts (long: treatment 93.76 to 93.87, control 91.87 to 92.72 tok/s). The prefill gain is not proven: on the long prompt the treatment range (4277 to 4440 tok/s) is inside the control range (4117 to 4502 tok/s).
Cold first request
The first measured request of the treatment arm in round 1 was slow on the short prompt (TTFT 1.09 s on Gemma, 0.71 s on Qwen 3.5; about 0.41 s in the other runs). This occurred on the first use of the new Metal kernel library. It did not occur in round 2. The medians are not changed by this run.
Output
The output text hash is the same in all 12 runs of each prompt and each model, in the two arms.
Where the gain comes from
| Phase | Code path | Change in ml-explore#4572 |
|---|---|---|
| Prefill | gather_qmm_rhs_nax (sorted route, NAX) |
Row tiles are scheduled from per-expert offsets. One tile does one matmul pass. |
| Decode | gather_qmv (8 assignments per token) |
No kernel change. ensure_row_contiguous_matrix returns at once for a row-contiguous input, which removes work from each call. |
Port notes
The cherry-pick of 90089440 had 9 conflicts in 4 files. The conflicted kernel functions were replaced with the upstream versions, with two adaptations for the fork base: the floating-point kernels have no global-scale input, and the NAX kernels keep the fork bound for the last K tile. The affine kernels, which these two models use, are identical to upstream.
Summary
Backport of two upstream commits onto fork
main.a2a09fd5gather_mm: row tiles are scheduled from per-expert offsets. New kernelgather_mm_offsets.90089440gather_qmm: the same scheduling in the sorted quantized kernels.ensure_row_contiguous_matrixreturns at once for a row-contiguous input.fork.yamlhas a section that describes the backport.scripts/check_forkdiff.pypasses.Port notes
Layers
Test
Benchmark: Qwen 3.5 35B A3B and Gemma 4 26B A4B, control against this stack
Result: decode is 1.3% to 1.4% faster on the two models. Prefill is 3.4% to 6.5% faster on Qwen 3.5. Prefill on Gemma is 1.5% to 2.5% faster, which is inside the run-to-run spread. Peak memory does not change. Output text is identical.
Setup
mlx-community/gemma-4-26B-A4B-it-qat-4bit;EigenLabs/Qwen3.5-35B-A3B-MLX-VL-4bit-g64at89646534d679360(mlx3fa8f25e)ff1eb060(mlxaa7af3eb, mlx-swiftdee30fb): ml-explore#4567 and ml-explore#4572darkbloom start --local, MTP off, 1 concurrent request, contiguous KVPrefill rate is prompt tokens divided by TTFT (client clock). TTFT is the time to the first streamed delta, reasoning text included. Decode rate is from stream chunk times. Values are medians. Peak memory is the server process physical footprint in the measured runs (n = 2 server starts per arm).
Qwen 3.5 35B A3B
Read: all speed metrics improve, and the gains repeat in the two rounds. On the long prompt, each treatment run is faster than the control run in the same position (run 1: 3817 and 3827 against 3549 and 3594 tok/s; runs 2 and 3: 4129 to 4168 against 3870 to 3931 tok/s). On the short prompt, the decode ranges do not overlap (treatment 130.3 to 132.0, control 128.9 to 130.1 tok/s). The long-prompt memory difference comes from one control start (28.14 GiB); the other three starts are 27.10 to 27.12 GiB.
Gemma 4 26B A4B
Read: the decode gain is measurable. The decode ranges do not overlap on the two prompts (long: treatment 93.76 to 93.87, control 91.87 to 92.72 tok/s). The prefill gain is not proven: on the long prompt the treatment range (4277 to 4440 tok/s) is inside the control range (4117 to 4502 tok/s).
Cold first request
The first measured request of the treatment arm in round 1 was slow on the short prompt (TTFT 1.09 s on Gemma, 0.71 s on Qwen 3.5; about 0.41 s in the other runs). This occurred on the first use of the new Metal kernel library. It did not occur in round 2. The medians are not changed by this run.
Output
The output text hash is the same in all 12 runs of each prompt and each model, in the two arms.
Where the gain comes from
gather_qmm_rhs_nax(sorted route, NAX)gather_qmv(8 assignments per token)ensure_row_contiguous_matrixreturns at once for a row-contiguous input, which removes work from each call.🤖 Generated with Claude Code
Refactor Pass
A dedicated refactor subagent reviewed the diff against the Layr-Labs
AGENTS.mdstructure, clean-code and refactor rules (d-inference ml-explore#1338). Outcome: no change. This is a backport of upstream code, and a cleanup would only make it differ from upstream.patch-id --stableas upstream a2a09fd.global_scaleargument and the fork'sBK - kk1bound, andfork.yamllists both. Upstream has no later commits to the touched files.scripts/check_forkdiff.pypasses: 181 changed files, 181 described. The branch merges cleanly with main abd33ee. The consumer pins in mlx-swift#34 and d-inference#1237 stay valid.Confirmed runtime requirement:
gather_mm_offsets()(matmul.cpp:2145) loadsgather_mm_offsetsfrom the prebuilt metallib, with no JIT fallback. Every sortedgather_mmpath calls it, and so does every sortedgather_qmmpath except the Gemma4 expert-tile route. So mlx-swift#34's SwiftPMKERNEL_LISTmust includegather_mm_offsets.metal. Without it, the first sorted gather call fails withUnable to load kernel gather_mm_offsets.Related, outside this PR: the fork branch
fix/gather-qmm-rhs-stale-M(07896d9) fixes a stale-Mbug on the same sortedgather_qmmroute.