Skip to content

build: pin MLX to the gather_mm and gather_qmm row-tile backport (ml-explore/mlx #4567, #4572) - #1237

Draft
davidtai wants to merge 2 commits into
masterfrom
feat/gather-mm-row-tiles
Draft

davidtai wants to merge 2 commits into
masterfrom
feat/gather-mm-row-tiles

Conversation

@davidtai

@davidtai davidtai commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Moves two submodule pins to the branches that carry upstream ml-explore/mlx#4567 (gather_mm) and ml-explore/mlx#4572 (gather_qmm).

Submodule From To
libs/mlx 3fa8f25e aa7af3eb (Layr-Labs/mlx#27)
libs/mlx-swift 0f4fe403 dee30fbe (Layr-Labs/mlx-swift#34)

Effect

Quantized MoE runners use gather_qmm. On an M5 Max, against master 4d679360:

Model Prefill Decode Peak memory
Qwen 3.5 35B A3B +3.4% to +6.5% +1.4% to +1.5% no change
Gemma 4 26B A4B +2.2% to +2.5% (inside the run-to-run spread) +1.3% no change

Output text is identical. The full tables are in the section below.

Test

  • Release build of darkbloom and the source-matched Metal kernel library: pass.
  • Benchmark on Qwen 3.5 35B A3B and Gemma 4 26B A4B: see the section below.
  • The provider unit tests were not run.

Draft. The pins point to branches that are not merged. Before merge, move the pins to the merged commits and update the pin table in docs/developer/build.md.

Benchmark: Qwen 3.5 35B A3B and Gemma 4 26B A4B, control against this stack

Result: decode is 1.3% to 1.4% faster on the two models. Prefill is 3.4% to 6.5% faster on Qwen 3.5. Prefill on Gemma is 1.5% to 2.5% faster, which is inside the run-to-run spread. Peak memory does not change. Output text is identical.

Setup

Item Value
Machine Apple M5 Max, 128 GiB, macOS 26.6.2
Models mlx-community/gemma-4-26B-A4B-it-qat-4bit; EigenLabs/Qwen3.5-35B-A3B-MLX-VL-4bit-g64 at 8964653
Control d-inference 4d679360 (mlx 3fa8f25e)
Treatment d-inference ff1eb060 (mlx aa7af3eb, mlx-swift dee30fb): #4567 and #4572
Server darkbloom start --local, MTP off, 1 concurrent request, contiguous KV
Sampling greedy, seed 20260928, max 1,024 output tokens
Design 2 interleaved rounds (control, treatment), 1 server start per arm per round, 3 measured runs per prompt. n = 6 per cell.
Conditions GPU lock held, resident service unloaded, one window for the two models

Prefill rate is prompt tokens divided by TTFT (client clock). TTFT is the time to the first streamed delta, reasoning text included. Decode rate is from stream chunk times. Values are medians. Peak memory is the server process physical footprint in the measured runs (n = 2 server starts per arm).

Qwen 3.5 35B A3B

Prompt Metric Control Treatment Δ treatment − control
Long (14,666 tokens) Prefill (tok/s) 3891 4143 +6.49%
Long Decode (tok/s) 117.4 119.1 +1.42%
Long TTFT (s) 3.770 3.540 −6.09%
Long Wall time (s) 12.50 12.13 −2.95%
Long Peak memory (GiB) 27.63 27.11 −1.88%
Short (1,886 tokens) Prefill (tok/s) 4570 4726 +3.41%
Short Decode (tok/s) 129.5 131.4 +1.50%
Short TTFT (s) 0.4127 0.3990 −3.32%
Short Wall time (s) 8.312 8.180 −1.59%
Short Peak memory (GiB) 21.96 21.96 −0.00%

Read: all speed metrics improve, and the gains repeat in the two rounds. On the long prompt, each treatment run is faster than the control run in the same position (run 1: 3817 and 3827 against 3549 and 3594 tok/s; runs 2 and 3: 4129 to 4168 against 3870 to 3931 tok/s). On the short prompt, the decode ranges do not overlap (treatment 130.3 to 132.0, control 128.9 to 130.1 tok/s). The long-prompt memory difference comes from one control start (28.14 GiB); the other three starts are 27.10 to 27.12 GiB.

Gemma 4 26B A4B

Prompt Metric Control Treatment Δ treatment − control
Long (16,383 tokens) Prefill (tok/s) 4299 4393 +2.20%
Long Decode (tok/s) 92.60 93.82 +1.32%
Long TTFT (s) 3.813 3.729 −2.19%
Long Wall time (s) 14.87 14.64 −1.57%
Long Peak memory (GiB) 22.46 22.46 +0.04%
Short (2,059 tokens) Prefill (tok/s) 4921 5045 +2.52%
Short Decode (tok/s) 108.1 109.4 +1.26%
Short TTFT (s) 0.4184 0.4082 −2.44%
Short Wall time (s) 6.345 6.274 −1.12%
Short Peak memory (GiB) 17.05 17.06 +0.05%

Read: the decode gain is measurable. The decode ranges do not overlap on the two prompts (long: treatment 93.76 to 93.87, control 91.87 to 92.72 tok/s). The prefill gain is not proven: on the long prompt the treatment range (4277 to 4440 tok/s) is inside the control range (4117 to 4502 tok/s).

Cold first request

The first measured request of the treatment arm in round 1 was slow on the short prompt (TTFT 1.09 s on Gemma, 0.71 s on Qwen 3.5; about 0.41 s in the other runs). This occurred on the first use of the new Metal kernel library. It did not occur in round 2. The medians are not changed by this run.

Output

The output text hash is the same in all 12 runs of each prompt and each model, in the two arms.

Where the gain comes from

Phase Code path Change in #4572
Prefill gather_qmm_rhs_nax (sorted route, NAX) Row tiles are scheduled from per-expert offsets. One tile does one matmul pass.
Decode gather_qmv (8 assignments per token) No kernel change. ensure_row_contiguous_matrix returns at once for a row-contiguous input, which removes work from each call.

🤖 Generated with Claude Code

libs/mlx -> Layr-Labs/mlx feat/gather-mm-row-tiles
libs/mlx-swift -> Layr-Labs/mlx-swift feat/gather-mm-row-tiles

Co-Authored-By: Claude Fable 5.1 <[email protected]>
@vercel

vercel Bot commented Sep 28, 2026

Copy link
Copy Markdown

Deployment failed for project d-inference with the following error:

You don't have permission to create a Preview Deployment for this Vercel project: d-inference.

View Documentation: https://vercel.com/docs/accounts/team-members-and-roles

@davidtai

davidtai commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

Superseded. This run measured a stack with ml-explore/mlx #4567 only. #4567 does not change gather_qmm. See the later comment "Benchmark with the gather_qmm change (ml-explore/mlx #4572)" for the result of the current stack.

Benchmark: Gemma 4 26B A4B runner, control against this stack

Result: no measurable change. This is the expected result. The runner uses gather_qmm, and this PR changes gather_mm only.

Setup

Item Value
Machine Apple M5 Max, 128 GiB, macOS 26.6.2
Model mlx-community/gemma-4-26B-A4B-it-qat-4bit
Control d-inference 4d679360 (mlx 3fa8f25e)
Treatment d-inference be166fe1 (mlx 11bf7814, mlx-swift b07f6b6a)
Server darkbloom start --local, MTP off, 1 concurrent request, contiguous KV
Sampling greedy, seed 20260928, max 1,024 output tokens
Design 2 interleaved rounds (control, treatment), 1 server start per arm per round, 3 measured runs per prompt. n = 6 per cell.
Conditions GPU lock held, resident service unloaded, 113.3 GiB available at start

Prefill rate is prompt tokens divided by TTFT (client clock). Decode rate is from stream chunk times. Peak memory is the server process physical footprint (n = 2 server starts per arm).

Long prompt (16,383 prompt tokens, 1,024 output tokens)

Metric Control Treatment Δ treatment − control
Prefill (tok/s) 4280 4378 +2.31%
Decode (tok/s) 92.57 92.75 +0.19%
TTFT (s) 3.836 3.744 −2.41%
Wall time (s) 14.89 14.77 −0.77%
Peak memory (GiB) 21.96 22.49 +2.45%

Read: the differences are smaller than the spread of the control arm (prefill 3943 to 4511 tok/s; peak memory 21.43 and 22.49 GiB across its two starts). In round 2 the arms agree to 0.4% on prefill and to 0.001 GiB on memory.

Short prompt (2,059 prompt tokens, 642 output tokens)

Metric Control Treatment Δ treatment − control
Prefill (tok/s, median) 4980 4965 −0.30%
Decode (tok/s) 108.2 108.2 −0.00%
TTFT (s, median) 0.4135 0.4147 +0.29%
Wall time (s, median) 6.343 6.341 −0.03%
Peak memory (GiB) 17.04 17.05 +0.05%

Read: no change. Medians are given because the first request of round 1 was slow in the two arms (TTFT 1.19 s and 1.22 s).

Output

The output text hash is the same in all 12 runs of each prompt, in the two arms.

Limit of this benchmark

This benchmark shows that the pin move does not change a gather_qmm runner. It does not measure the speedup of the PR. That measurement needs a model with unquantized experts.

@davidtai

davidtai commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

Superseded. This run measured a stack with ml-explore/mlx #4567 only. #4567 does not change gather_qmm. See the later comment "Benchmark with the gather_qmm change (ml-explore/mlx #4572)" for the result of the current stack.

Benchmark: Qwen 3.5 35B A3B runner, control against this stack

Result: no measurable change. This is the expected result. The runner uses gather_qmm, and the upstream change applies to gather_mm only.

Setup

Item Value
Machine Apple M5 Max, 128 GiB, macOS 26.6.2
Model EigenLabs/Qwen3.5-35B-A3B-MLX-VL-4bit-g64 at 8964653 (catalog model qwen3.5-35b-a3b)
Control d-inference 4d679360 (mlx 3fa8f25e)
Treatment d-inference be166fe1 (mlx 11bf7814, mlx-swift b07f6b6a)
Server darkbloom start --local, MTP off, 1 concurrent request, contiguous KV
Sampling greedy, seed 20260928, max 1,024 output tokens
Design 2 interleaved rounds (control, treatment), 1 server start per arm per round, 3 measured runs per prompt. n = 6 per cell.
Conditions GPU lock held, resident service unloaded, 113.4 GiB available at start

The builds and the prompts are the same as in the Gemma benchmark. Prefill rate is prompt tokens divided by TTFT (client clock). TTFT is the time to the first streamed delta, reasoning text included. Decode rate is from stream chunk times. Values are medians, because a small number of runs were slow in each arm.

Long prompt (14,666 prompt tokens, 1,024 output tokens)

Metric Control Treatment Δ treatment − control
Prefill (tok/s) 3920 3806 −2.92%
Decode (tok/s) 117.6 116.9 −0.63%
TTFT (s) 3.741 3.857 +3.09%
Wall time (s) 12.44 12.63 +1.50%
Peak memory in the measured runs (GiB) 27.16 27.11 −0.21%

Read: the difference comes from round 1 only, where the treatment runs were slower (prefill 3501 to 3694 tok/s). In round 2 the arms agree to 0.3% on prefill (3915 and 3926 tok/s, mean) and to 0.6% on decode. The sign is opposite to the Gemma long-prompt result (+2.31%). The two results are consistent with run-to-run noise.

Short prompt (1,886 prompt tokens, 1,024 output tokens)

Metric Control Treatment Δ treatment − control
Prefill (tok/s) 4563 4555 −0.18%
Decode (tok/s) 129.9 128.9 −0.77%
TTFT (s) 0.4133 0.4140 +0.17%
Wall time (s) 8.325 8.356 +0.37%
Peak memory in the measured runs (GiB) 21.95 21.96 +0.03%

Read: no change in prefill, TTFT, or memory. The decode difference (0.8%) is of the same size as the difference between the two control rounds (130.3 and 129.1 tok/s, mean).

Process peak memory

The lifetime maximum of the server footprint is 30.94 GiB in the control arm and 30.95 GiB in the treatment arm (+0.02%). This maximum occurs before the measured runs (model load and warmup).

Output

The output text hash is the same in all 12 runs of each prompt, in the two arms.

Limit of this benchmark

This benchmark shows that the pin move does not change a gather_qmm runner. It does not measure the speedup of the upstream change. That measurement needs a model with unquantized experts.

@vercel

vercel Bot commented Sep 28, 2026

Copy link
Copy Markdown

Deployment failed for project eigen-homepages-darkbloom with the following error:

You don't have permission to create a Preview Deployment for this Vercel project: eigen-homepages-darkbloom.

View Documentation: https://vercel.com/docs/accounts/team-members-and-roles

@vercel

vercel Bot commented Sep 28, 2026

Copy link
Copy Markdown

Deployment failed for project d-inference-landing with the following error:

You don't have permission to create a Preview Deployment for this Vercel project: d-inference-landing.

View Documentation: https://vercel.com/docs/accounts/team-members-and-roles

@vercel

vercel Bot commented Sep 28, 2026

Copy link
Copy Markdown

Deployment failed for project d-inference-console-ui-dev with the following error:

You don't have permission to create a Preview Deployment for this Vercel project: d-inference-console-ui-dev.

View Documentation: https://vercel.com/docs/accounts/team-members-and-roles

@davidtai

Copy link
Copy Markdown
Contributor Author

Benchmark with the gather_qmm change (ml-explore/mlx #4572): Gemma 4 26B A4B and Qwen 3.5 35B A3B

Result: decode is 1.3% to 1.4% faster on the two models. Prefill is 3.4% to 6.5% faster on Qwen 3.5. Prefill on Gemma is 1.5% to 2.5% faster, which is inside the run-to-run spread. Peak memory does not change. Output text is identical.

The two earlier benchmark comments on this PR measured a stack with #4567 only, which does not change gather_qmm. This comment replaces them as the result for this stack.

Setup

Item Value
Machine Apple M5 Max, 128 GiB, macOS 26.6.2
Models mlx-community/gemma-4-26B-A4B-it-qat-4bit; EigenLabs/Qwen3.5-35B-A3B-MLX-VL-4bit-g64 at 8964653
Control d-inference 4d679360 (mlx 3fa8f25e)
Treatment d-inference ff1eb060 (mlx aa7af3eb, mlx-swift dee30fb): #4567 and #4572
Server darkbloom start --local, MTP off, 1 concurrent request, contiguous KV
Sampling greedy, seed 20260928, max 1,024 output tokens
Design 2 interleaved rounds (control, treatment), 1 server start per arm per round, 3 measured runs per prompt. n = 6 per cell.
Conditions GPU lock held, resident service unloaded, one window for the two models

Prefill rate is prompt tokens divided by TTFT (client clock). TTFT is the time to the first streamed delta, reasoning text included. Decode rate is from stream chunk times. Values are medians. Peak memory is the server process physical footprint in the measured runs (n = 2 server starts per arm).

Qwen 3.5 35B A3B

Prompt Metric Control Treatment Δ treatment − control
Long (14,666 tokens) Prefill (tok/s) 3891 4143 +6.49%
Long Decode (tok/s) 117.4 119.1 +1.42%
Long TTFT (s) 3.770 3.540 −6.09%
Long Wall time (s) 12.50 12.13 −2.95%
Long Peak memory (GiB) 27.63 27.11 −1.88%
Short (1,886 tokens) Prefill (tok/s) 4570 4726 +3.41%
Short Decode (tok/s) 129.5 131.4 +1.50%
Short TTFT (s) 0.4127 0.3990 −3.32%
Short Wall time (s) 8.312 8.180 −1.59%
Short Peak memory (GiB) 21.96 21.96 −0.00%

Read: all speed metrics improve, and the gains repeat in the two rounds. On the long prompt, each treatment run is faster than the control run in the same position (run 1: 3817 and 3827 against 3549 and 3594 tok/s; runs 2 and 3: 4129 to 4168 against 3870 to 3931 tok/s). On the short prompt, the decode ranges do not overlap (treatment 130.3 to 132.0, control 128.9 to 130.1 tok/s). The long-prompt memory difference comes from one control start (28.14 GiB); the other three starts are 27.10 to 27.12 GiB.

Gemma 4 26B A4B

Prompt Metric Control Treatment Δ treatment − control
Long (16,383 tokens) Prefill (tok/s) 4299 4393 +2.20%
Long Decode (tok/s) 92.60 93.82 +1.32%
Long TTFT (s) 3.813 3.729 −2.19%
Long Wall time (s) 14.87 14.64 −1.57%
Long Peak memory (GiB) 22.46 22.46 +0.04%
Short (2,059 tokens) Prefill (tok/s) 4921 5045 +2.52%
Short Decode (tok/s) 108.1 109.4 +1.26%
Short TTFT (s) 0.4184 0.4082 −2.44%
Short Wall time (s) 6.345 6.274 −1.12%
Short Peak memory (GiB) 17.05 17.06 +0.05%

Read: the decode gain is measurable. The decode ranges do not overlap on the two prompts (long: treatment 93.76 to 93.87, control 91.87 to 92.72 tok/s). The prefill gain is not proven: on the long prompt the treatment range (4277 to 4440 tok/s) is inside the control range (4117 to 4502 tok/s).

Cold first request

The first measured request of the treatment arm in round 1 was slow on the short prompt (TTFT 1.09 s on Gemma, 0.71 s on Qwen 3.5; about 0.41 s in the other runs). This occurred on the first use of the new Metal kernel library. It did not occur in round 2. The medians are not changed by this run.

Output

The output text hash is the same in all 12 runs of each prompt and each model, in the two arms.

Where the gain comes from

Phase Code path Change in #4572
Prefill gather_qmm_rhs_nax (sorted route, NAX) Row tiles are scheduled from per-expert offsets. One tile does one matmul pass.
Decode gather_qmv (8 assignments per token) No kernel change. ensure_row_contiguous_matrix returns at once for a row-contiguous input, which removes work from each call.

Port notes

The cherry-pick of 90089440 had 9 conflicts in 4 files. The conflicted kernel functions were replaced with the upstream versions, with two adaptations for the fork base: the floating-point kernels have no global-scale input, and the NAX kernels keep the fork bound for the last K tile. The affine kernels, which these two models use, are identical to upstream.

@davidtai davidtai changed the title build: pin MLX to the gather_mm row-tile backport (ml-explore/mlx #4567) build: pin MLX to the gather_mm and gather_qmm row-tile backport (ml-explore/mlx #4567, #4572) Sep 28, 2026

This branch is waiting to be deployed

1 waiting deployment
benchmarks — ff1eb060 Waiting Sep 28, 2026 by davidtai via E2E Benchmarks #1780
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant