Skip to content

First run after a local mlx build: prefill up to 2.3x slower on small models (Metal shader cache cold), negligible in absolute time #4566

Description

@jrp2014

This is an informational note, not a bug report. After a fresh source build of mlx, the first run shows slower prefill throughput. The effect is large as a ratio for small, fast models and negligible in wall time. I'm recording it because it can mislead anyone who compares throughput across builds.

Environment

What happened

The same 48-model sweep ran twice on the same image, prompt and settings, with no code or model changes between runs:

  • mlx built at 21:54
  • run 1 (cold): 22:06 to 22:19
  • run 2 (warm): 23:27 to 23:40

Generated text was identical for all 46 models that completed. Total prefill time was also the same, at 392 s in both runs. However, the smallest models prefilled up to 2.3x faster in run 2:

Model Prompt tokens Prefill, run 1 (s) Prefill, run 2 (s) Speed-up
gemma-4-e4b-it-4bit 592 0.43 0.19 2.30
granite-4.0-3b-vision-4bit 1383 0.90 0.39 2.28
gemma-3n-E4B-it-4bit 590 0.67 0.31 2.17
Molmo2-8B-4bit 1526 1.60 0.75 2.13
LFM2.5-VL-450M-MLX-bf16 2119 0.21 0.11 1.89
nanoLLaVA-1.5-4bit 332 0.17 0.09 1.85

Each of these models saved under one second. Larger models, with 30-60 s prefills, moved by about ±10-20% in either direction, which is ordinary run-to-run variation. Other work was running on the machine during part of run 2.

Why I think it is the Metal shader cache

Run 1 was the first process to use the newly built mlx.metallib. When it exited at about 22:22, the per-user Metal cache ($(getconf DARWIN_USER_CACHE_DIR)com.apple.metal) gained about 35 MB of functions.data and 44 MB of libraries.data. Run 2 wrote no new data there and only updated the index files. So run 1 compiled pipeline states on first use, and run 2 reused them. Within a single run, the cost repeats for every kernel variant a model is first to use. That is why models with sub-second prefills show it as a large ratio. I haven't isolated it further, for example by clearing the cache and rerunning one model.

Impact

  • Wall time is negligible, at a few seconds across a 13-minute sweep.
  • Throughput comparisons are affected. The first run after any rebuild, or after an OS/Xcode update that invalidates the cache, reports inflated prefill times for fast models. A 2x "regression" on a sub-second prefill can be just this.

Possible options (no action needed)

  • A note in the build/benchmarking docs: do one warm-up run after building from source before timing prefill.
  • Longer term, pipeline precompilation such as MTLBinaryArchive at build or install time would remove the effect. I realise that may not be worth its complexity for a sub-second cost.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions