Research preview: reproducible MoE routing, cache-policy, and memory-oversubscription experiments for consumer NVIDIA systems.
中文说明 · Reproduction guide · Published evidence · Contributing
MoE Expert Cache is a llama.cpp companion research lab, not a new general-purpose inference engine. It keeps a pinned llama.cpp revision, a small opt-in patch series, provenance-aware launch and benchmark tools, deterministic routing/cache analysis, and an explicit record of failed ideas. No model weights are distributed.
This repository is an early research preview. It does not claim a universal cache speedup, validated 122B throughput, or support for arbitrary MoE checkpoints.
| Finding | Controlled evidence | Interpretation |
|---|---|---|
| Demand-mapped Qwen3.5-35B-A3B under an exact 8 GiB RAM cgroup | At 128 generated tokens, median wall time fell from 828.3091 s to 100.9432 s and decode rose from 0.1595 to 1.8929 token/s; matched-length raw output was byte-identical | Useful only when the model genuinely exceeds controlled RAM plus observed VRAM; prefill regressed |
| Earlier Qwen3-30B-A3B oversubscription gate | At 16 tokens, median wall time fell from 90.096 s to 33.175 s and decode rose from 0.23 to 1.99 token/s | Mechanism confirmation on a different GGUF, not Qwen3.5 performance |
| Physical GPU-hot / CPU-miss expert cache | 49.3% median hit rate, but 17.305 token/s versus 20.795 for the fair CPU-MoE baseline and 26.015 for automatic placement | Rejected for both performance and deterministic-output failure |
| Qwen3.5-122B-A10B on a proposed 64 GiB RAM / 32 GB RTX 5090 D system | Static aggregate byte planning only | No 122B weights were downloaded; placement, speed, quality, and stability remain unverified |
Sanitized, machine-readable summaries and digests of the retained canonical
results are in evidence/.
- CPU-backend Qwen3-MoE route tracing at the llama.cpp
MUL_MAT_IDboundary - deterministic route replay through the Least-Stale slot policy
- opt-in Linux demand mmap for severe host-memory oversubscription
- conservative Qwen3.5 startup planning from model metadata and real memory constraints
- provenance binding for the llama.cpp binary, runtime libraries, patches, model manifest, and benchmark schedule
- explicit separation of active, experimental, and rejected patches
The active series does not implement a production GPU expert cache.
Python 3.12 is required. This path downloads no model and needs no CUDA:
python3 -m venv .venv-ci
source .venv-ci/bin/activate
python -m pip install -r requirements-ci.txt
python scripts/check_public_release.py --bootstrap
python -m pytest \
tests/test_llamacpp_moe_trace.py \
tests/test_llamacpp_oversubscription.py \
tests/test_task86_qwen35_benchmark.py \
tests/test_task86_llamacpp_provenance.py--bootstrap is needed only before the first curated Git commit. CI uses the
strict tracked-file check. See docs/reproducing.md
for the patch build and hardware-gated levels.
The public preview targets exactly:
https://github.com/ggml-org/llama.cpp.git
b2f221684fcd898e947a121baeda80f345da3e6b
Apply the active series:
git clone https://github.com/ggml-org/llama.cpp.git
git -C llama.cpp checkout --detach b2f221684fcd898e947a121baeda80f345da3e6b
bash scripts/apply_llamacpp_fork_patches.sh "$PWD/llama.cpp"engine/llama.cpp.lock.json is the source of
truth. The active series ends at 0004-demand-mmap-oversubscription.patch.
Patches 0005 through 0008 and everything under rejected/ are retained for
review but never applied by the normal script.
- Qwen3.5-35B physical results apply only to the exact 17,821,719,168-byte
UD-IQ4_NL GGUF with SHA-256
6aec31ed654fbf12f9950c98bab0a09faea433ffdc4b3acfeb55a1f050e3e444. - The 8 GiB cgroup is a controlled proxy; the WSL host retained its normal physical memory.
- Whole-device GPU memory was observed, not capped.
- The benchmark did not capture usable cgroup
io.statbytes, so it does not establish physical SSD traffic. - Demand mmap is an oversubscribed long-decode mode, not a default for systems with sufficient usable RAM.
- The retained local llama.cpp binary was built for SM89 and cannot be reused on SM120 hardware such as RTX 5090-class GPUs.
engine/— pinned upstream identity and active patch lockpatches/llama.cpp/— active, experimental, and rejected diffssrc/moe_cache/— public trace, policy, and oversubscription helpersscripts/— patch, launch, provenance, benchmark, and release checkstests/— no-model correctness and release gatesevidence/— sanitized summaries, never model weights or raw outputdocs/— reproduction, publication scope, and Chinese overview
The first public commit is a curated surface, not a dump of the private
workspace. See docs/publication-scope.md.
llama.cpp remains the product-oriented downstream engine and the fair baseline. The project contributes measurement, provenance, focused patches, and negative evidence around memory-constrained MoE inference. A native executor remains a possible future research branch, not an implemented feature.
No model is bundled or automatically downloaded. Each performance claim is scoped to an exact checkpoint or GGUF, revision, quantization, and hash. Model weights, raw traces, generated work directories, credentials, private network locations, and machine-specific storage records must not be committed.
- Contributing
- Security
- Support
- Code of Conduct
- Third-party notices
- Citation metadata
- Known limitations
Project-authored code and documentation are available under the MIT License. Third-party code represented by the llama.cpp patch series remains subject to the upstream copyright and license notices described in THIRD_PARTY_NOTICES.md.