Conversation
libs/mlx -> Layr-Labs/mlx feat/gather-mm-row-tiles libs/mlx-swift -> Layr-Labs/mlx-swift feat/gather-mm-row-tiles Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
Deployment failed for project d-inference with the following error: View Documentation: https://vercel.com/docs/accounts/team-members-and-roles |
Benchmark: Gemma 4 26B A4B runner, control against this stackResult: no measurable change. This is the expected result. The runner uses Setup
Prefill rate is prompt tokens divided by TTFT (client clock). Decode rate is from stream chunk times. Peak memory is the server process physical footprint (n = 2 server starts per arm). Long prompt (16,383 prompt tokens, 1,024 output tokens)
Read: the differences are smaller than the spread of the control arm (prefill 3943 to 4511 tok/s; peak memory 21.43 and 22.49 GiB across its two starts). In round 2 the arms agree to 0.4% on prefill and to 0.001 GiB on memory. Short prompt (2,059 prompt tokens, 642 output tokens)
Read: no change. Medians are given because the first request of round 1 was slow in the two arms (TTFT 1.19 s and 1.22 s). OutputThe output text hash is the same in all 12 runs of each prompt, in the two arms. Limit of this benchmarkThis benchmark shows that the pin move does not change a |
Benchmark: Qwen 3.5 35B A3B runner, control against this stackResult: no measurable change. This is the expected result. The runner uses Setup
The builds and the prompts are the same as in the Gemma benchmark. Prefill rate is prompt tokens divided by TTFT (client clock). TTFT is the time to the first streamed delta, reasoning text included. Decode rate is from stream chunk times. Values are medians, because a small number of runs were slow in each arm. Long prompt (14,666 prompt tokens, 1,024 output tokens)
Read: the difference comes from round 1 only, where the treatment runs were slower (prefill 3501 to 3694 tok/s). In round 2 the arms agree to 0.3% on prefill (3915 and 3926 tok/s, mean) and to 0.6% on decode. The sign is opposite to the Gemma long-prompt result (+2.31%). The two results are consistent with run-to-run noise. Short prompt (1,886 prompt tokens, 1,024 output tokens)
Read: no change in prefill, TTFT, or memory. The decode difference (0.8%) is of the same size as the difference between the two control rounds (130.3 and 129.1 tok/s, mean). Process peak memoryThe lifetime maximum of the server footprint is 30.94 GiB in the control arm and 30.95 GiB in the treatment arm (+0.02%). This maximum occurs before the measured runs (model load and warmup). OutputThe output text hash is the same in all 12 runs of each prompt, in the two arms. Limit of this benchmarkThis benchmark shows that the pin move does not change a |
…572) Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
Deployment failed for project eigen-homepages-darkbloom with the following error: View Documentation: https://vercel.com/docs/accounts/team-members-and-roles |
|
Deployment failed for project d-inference-landing with the following error: View Documentation: https://vercel.com/docs/accounts/team-members-and-roles |
|
Deployment failed for project d-inference-console-ui-dev with the following error: View Documentation: https://vercel.com/docs/accounts/team-members-and-roles |
Benchmark with the
|
| Item | Value |
|---|---|
| Machine | Apple M5 Max, 128 GiB, macOS 26.6.2 |
| Models | mlx-community/gemma-4-26B-A4B-it-qat-4bit; EigenLabs/Qwen3.5-35B-A3B-MLX-VL-4bit-g64 at 8964653 |
| Control | d-inference 4d679360 (mlx 3fa8f25e) |
| Treatment | d-inference ff1eb060 (mlx aa7af3eb, mlx-swift dee30fb): #4567 and #4572 |
| Server | darkbloom start --local, MTP off, 1 concurrent request, contiguous KV |
| Sampling | greedy, seed 20260928, max 1,024 output tokens |
| Design | 2 interleaved rounds (control, treatment), 1 server start per arm per round, 3 measured runs per prompt. n = 6 per cell. |
| Conditions | GPU lock held, resident service unloaded, one window for the two models |
Prefill rate is prompt tokens divided by TTFT (client clock). TTFT is the time to the first streamed delta, reasoning text included. Decode rate is from stream chunk times. Values are medians. Peak memory is the server process physical footprint in the measured runs (n = 2 server starts per arm).
Qwen 3.5 35B A3B
| Prompt | Metric | Control | Treatment | Δ treatment − control |
|---|---|---|---|---|
| Long (14,666 tokens) | Prefill (tok/s) | 3891 | 4143 | +6.49% |
| Long | Decode (tok/s) | 117.4 | 119.1 | +1.42% |
| Long | TTFT (s) | 3.770 | 3.540 | −6.09% |
| Long | Wall time (s) | 12.50 | 12.13 | −2.95% |
| Long | Peak memory (GiB) | 27.63 | 27.11 | −1.88% |
| Short (1,886 tokens) | Prefill (tok/s) | 4570 | 4726 | +3.41% |
| Short | Decode (tok/s) | 129.5 | 131.4 | +1.50% |
| Short | TTFT (s) | 0.4127 | 0.3990 | −3.32% |
| Short | Wall time (s) | 8.312 | 8.180 | −1.59% |
| Short | Peak memory (GiB) | 21.96 | 21.96 | −0.00% |
Read: all speed metrics improve, and the gains repeat in the two rounds. On the long prompt, each treatment run is faster than the control run in the same position (run 1: 3817 and 3827 against 3549 and 3594 tok/s; runs 2 and 3: 4129 to 4168 against 3870 to 3931 tok/s). On the short prompt, the decode ranges do not overlap (treatment 130.3 to 132.0, control 128.9 to 130.1 tok/s). The long-prompt memory difference comes from one control start (28.14 GiB); the other three starts are 27.10 to 27.12 GiB.
Gemma 4 26B A4B
| Prompt | Metric | Control | Treatment | Δ treatment − control |
|---|---|---|---|---|
| Long (16,383 tokens) | Prefill (tok/s) | 4299 | 4393 | +2.20% |
| Long | Decode (tok/s) | 92.60 | 93.82 | +1.32% |
| Long | TTFT (s) | 3.813 | 3.729 | −2.19% |
| Long | Wall time (s) | 14.87 | 14.64 | −1.57% |
| Long | Peak memory (GiB) | 22.46 | 22.46 | +0.04% |
| Short (2,059 tokens) | Prefill (tok/s) | 4921 | 5045 | +2.52% |
| Short | Decode (tok/s) | 108.1 | 109.4 | +1.26% |
| Short | TTFT (s) | 0.4184 | 0.4082 | −2.44% |
| Short | Wall time (s) | 6.345 | 6.274 | −1.12% |
| Short | Peak memory (GiB) | 17.05 | 17.06 | +0.05% |
Read: the decode gain is measurable. The decode ranges do not overlap on the two prompts (long: treatment 93.76 to 93.87, control 91.87 to 92.72 tok/s). The prefill gain is not proven: on the long prompt the treatment range (4277 to 4440 tok/s) is inside the control range (4117 to 4502 tok/s).
Cold first request
The first measured request of the treatment arm in round 1 was slow on the short prompt (TTFT 1.09 s on Gemma, 0.71 s on Qwen 3.5; about 0.41 s in the other runs). This occurred on the first use of the new Metal kernel library. It did not occur in round 2. The medians are not changed by this run.
Output
The output text hash is the same in all 12 runs of each prompt and each model, in the two arms.
Where the gain comes from
| Phase | Code path | Change in #4572 |
|---|---|---|
| Prefill | gather_qmm_rhs_nax (sorted route, NAX) |
Row tiles are scheduled from per-expert offsets. One tile does one matmul pass. |
| Decode | gather_qmv (8 assignments per token) |
No kernel change. ensure_row_contiguous_matrix returns at once for a row-contiguous input, which removes work from each call. |
Port notes
The cherry-pick of 90089440 had 9 conflicts in 4 files. The conflicted kernel functions were replaced with the upstream versions, with two adaptations for the fork base: the floating-point kernels have no global-scale input, and the NAX kernels keep the fork bound for the last K tile. The affine kernels, which these two models use, are identical to upstream.
Summary
Moves two submodule pins to the branches that carry upstream ml-explore/mlx#4567 (
gather_mm) and ml-explore/mlx#4572 (gather_qmm).libs/mlx3fa8f25eaa7af3eb(Layr-Labs/mlx#27)libs/mlx-swift0f4fe403dee30fbe(Layr-Labs/mlx-swift#34)Effect
Quantized MoE runners use
gather_qmm. On an M5 Max, against master4d679360:Output text is identical. The full tables are in the section below.
Test
darkbloomand the source-matched Metal kernel library: pass.Draft. The pins point to branches that are not merged. Before merge, move the pins to the merged commits and update the pin table in
docs/developer/build.md.Benchmark: Qwen 3.5 35B A3B and Gemma 4 26B A4B, control against this stack
Result: decode is 1.3% to 1.4% faster on the two models. Prefill is 3.4% to 6.5% faster on Qwen 3.5. Prefill on Gemma is 1.5% to 2.5% faster, which is inside the run-to-run spread. Peak memory does not change. Output text is identical.
Setup
mlx-community/gemma-4-26B-A4B-it-qat-4bit;EigenLabs/Qwen3.5-35B-A3B-MLX-VL-4bit-g64at89646534d679360(mlx3fa8f25e)ff1eb060(mlxaa7af3eb, mlx-swiftdee30fb): #4567 and #4572darkbloom start --local, MTP off, 1 concurrent request, contiguous KVPrefill rate is prompt tokens divided by TTFT (client clock). TTFT is the time to the first streamed delta, reasoning text included. Decode rate is from stream chunk times. Values are medians. Peak memory is the server process physical footprint in the measured runs (n = 2 server starts per arm).
Qwen 3.5 35B A3B
Read: all speed metrics improve, and the gains repeat in the two rounds. On the long prompt, each treatment run is faster than the control run in the same position (run 1: 3817 and 3827 against 3549 and 3594 tok/s; runs 2 and 3: 4129 to 4168 against 3870 to 3931 tok/s). On the short prompt, the decode ranges do not overlap (treatment 130.3 to 132.0, control 128.9 to 130.1 tok/s). The long-prompt memory difference comes from one control start (28.14 GiB); the other three starts are 27.10 to 27.12 GiB.
Gemma 4 26B A4B
Read: the decode gain is measurable. The decode ranges do not overlap on the two prompts (long: treatment 93.76 to 93.87, control 91.87 to 92.72 tok/s). The prefill gain is not proven: on the long prompt the treatment range (4277 to 4440 tok/s) is inside the control range (4117 to 4502 tok/s).
Cold first request
The first measured request of the treatment arm in round 1 was slow on the short prompt (TTFT 1.09 s on Gemma, 0.71 s on Qwen 3.5; about 0.41 s in the other runs). This occurred on the first use of the new Metal kernel library. It did not occur in round 2. The medians are not changed by this run.
Output
The output text hash is the same in all 12 runs of each prompt and each model, in the two arms.
Where the gain comes from
gather_qmm_rhs_nax(sorted route, NAX)gather_qmv(8 assignments per token)ensure_row_contiguous_matrixreturns at once for a row-contiguous input, which removes work from each call.🤖 Generated with Claude Code