This is an informational note, not a bug report. After a fresh source build of mlx, the first run shows slower prefill throughput. The effect is large as a ratio for small, fast models and negligible in wall time. I'm recording it because it can mislead anyone who compares throughput across builds.
Environment
What happened
The same 48-model sweep ran twice on the same image, prompt and settings, with no code or model changes between runs:
- mlx built at 21:54
- run 1 (cold): 22:06 to 22:19
- run 2 (warm): 23:27 to 23:40
Generated text was identical for all 46 models that completed. Total prefill time was also the same, at 392 s in both runs. However, the smallest models prefilled up to 2.3x faster in run 2:
| Model |
Prompt tokens |
Prefill, run 1 (s) |
Prefill, run 2 (s) |
Speed-up |
| gemma-4-e4b-it-4bit |
592 |
0.43 |
0.19 |
2.30 |
| granite-4.0-3b-vision-4bit |
1383 |
0.90 |
0.39 |
2.28 |
| gemma-3n-E4B-it-4bit |
590 |
0.67 |
0.31 |
2.17 |
| Molmo2-8B-4bit |
1526 |
1.60 |
0.75 |
2.13 |
| LFM2.5-VL-450M-MLX-bf16 |
2119 |
0.21 |
0.11 |
1.89 |
| nanoLLaVA-1.5-4bit |
332 |
0.17 |
0.09 |
1.85 |
Each of these models saved under one second. Larger models, with 30-60 s prefills, moved by about ±10-20% in either direction, which is ordinary run-to-run variation. Other work was running on the machine during part of run 2.
Why I think it is the Metal shader cache
Run 1 was the first process to use the newly built mlx.metallib. When it exited at about 22:22, the per-user Metal cache ($(getconf DARWIN_USER_CACHE_DIR)com.apple.metal) gained about 35 MB of functions.data and 44 MB of libraries.data. Run 2 wrote no new data there and only updated the index files. So run 1 compiled pipeline states on first use, and run 2 reused them. Within a single run, the cost repeats for every kernel variant a model is first to use. That is why models with sub-second prefills show it as a large ratio. I haven't isolated it further, for example by clearing the cache and rerunning one model.
Impact
- Wall time is negligible, at a few seconds across a 13-minute sweep.
- Throughput comparisons are affected. The first run after any rebuild, or after an OS/Xcode update that invalidates the cache, reports inflated prefill times for fast models. A 2x "regression" on a sub-second prefill can be just this.
Possible options (no action needed)
- A note in the build/benchmarking docs: do one warm-up run after building from source before timing prefill.
- Longer term, pipeline precompilation such as
MTLBinaryArchive at build or install time would remove the effect. I realise that may not be worth its complexity for a sub-second cost.
This is an informational note, not a bug report. After a fresh source build of mlx, the first run shows slower prefill throughput. The effect is large as a ratio for small, fast models and negligible in wall time. I'm recording it because it can mislead anyone who compares throughput across builds.
Environment
pip install -e ".[dev]", plus the localgated_delta_update_nax.hchange needed to compile under Xcode 27 (seemlxno longer builds from source with Xcode 27 (macOS 27.0 SDK) #4533)What happened
The same 48-model sweep ran twice on the same image, prompt and settings, with no code or model changes between runs:
Generated text was identical for all 46 models that completed. Total prefill time was also the same, at 392 s in both runs. However, the smallest models prefilled up to 2.3x faster in run 2:
Each of these models saved under one second. Larger models, with 30-60 s prefills, moved by about ±10-20% in either direction, which is ordinary run-to-run variation. Other work was running on the machine during part of run 2.
Why I think it is the Metal shader cache
Run 1 was the first process to use the newly built
mlx.metallib. When it exited at about 22:22, the per-user Metal cache ($(getconf DARWIN_USER_CACHE_DIR)com.apple.metal) gained about 35 MB offunctions.dataand 44 MB oflibraries.data. Run 2 wrote no new data there and only updated the index files. So run 1 compiled pipeline states on first use, and run 2 reused them. Within a single run, the cost repeats for every kernel variant a model is first to use. That is why models with sub-second prefills show it as a large ratio. I haven't isolated it further, for example by clearing the cache and rerunning one model.Impact
Possible options (no action needed)
MTLBinaryArchiveat build or install time would remove the effect. I realise that may not be worth its complexity for a sub-second cost.