Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MoE Expert Cache

Research preview: reproducible MoE routing, cache-policy, and memory-oversubscription experiments for consumer NVIDIA systems.

中文说明 · Reproduction guide · Published evidence · Contributing

MoE Expert Cache is a llama.cpp companion research lab, not a new general-purpose inference engine. It keeps a pinned llama.cpp revision, a small opt-in patch series, provenance-aware launch and benchmark tools, deterministic routing/cache analysis, and an explicit record of failed ideas. No model weights are distributed.

This repository is an early research preview. It does not claim a universal cache speedup, validated 122B throughput, or support for arbitrary MoE checkpoints.

What Is Verified

Finding Controlled evidence Interpretation
Demand-mapped Qwen3.5-35B-A3B under an exact 8 GiB RAM cgroup At 128 generated tokens, median wall time fell from 828.3091 s to 100.9432 s and decode rose from 0.1595 to 1.8929 token/s; matched-length raw output was byte-identical Useful only when the model genuinely exceeds controlled RAM plus observed VRAM; prefill regressed
Earlier Qwen3-30B-A3B oversubscription gate At 16 tokens, median wall time fell from 90.096 s to 33.175 s and decode rose from 0.23 to 1.99 token/s Mechanism confirmation on a different GGUF, not Qwen3.5 performance
Physical GPU-hot / CPU-miss expert cache 49.3% median hit rate, but 17.305 token/s versus 20.795 for the fair CPU-MoE baseline and 26.015 for automatic placement Rejected for both performance and deterministic-output failure
Qwen3.5-122B-A10B on a proposed 64 GiB RAM / 32 GB RTX 5090 D system Static aggregate byte planning only No 122B weights were downloaded; placement, speed, quality, and stability remain unverified

Sanitized, machine-readable summaries and digests of the retained canonical results are in evidence/.

Included Capabilities

  • CPU-backend Qwen3-MoE route tracing at the llama.cpp MUL_MAT_ID boundary
  • deterministic route replay through the Least-Stale slot policy
  • opt-in Linux demand mmap for severe host-memory oversubscription
  • conservative Qwen3.5 startup planning from model metadata and real memory constraints
  • provenance binding for the llama.cpp binary, runtime libraries, patches, model manifest, and benchmark schedule
  • explicit separation of active, experimental, and rejected patches

The active series does not implement a production GPU expert cache.

Fifteen-Minute, No-Model Check

Python 3.12 is required. This path downloads no model and needs no CUDA:

python3 -m venv .venv-ci
source .venv-ci/bin/activate
python -m pip install -r requirements-ci.txt
python scripts/check_public_release.py --bootstrap
python -m pytest \
  tests/test_llamacpp_moe_trace.py \
  tests/test_llamacpp_oversubscription.py \
  tests/test_task86_qwen35_benchmark.py \
  tests/test_task86_llamacpp_provenance.py

--bootstrap is needed only before the first curated Git commit. CI uses the strict tracked-file check. See docs/reproducing.md for the patch build and hardware-gated levels.

llama.cpp Patch Series

The public preview targets exactly:

https://github.com/ggml-org/llama.cpp.git
b2f221684fcd898e947a121baeda80f345da3e6b

Apply the active series:

git clone https://github.com/ggml-org/llama.cpp.git
git -C llama.cpp checkout --detach b2f221684fcd898e947a121baeda80f345da3e6b
bash scripts/apply_llamacpp_fork_patches.sh "$PWD/llama.cpp"

engine/llama.cpp.lock.json is the source of truth. The active series ends at 0004-demand-mmap-oversubscription.patch. Patches 0005 through 0008 and everything under rejected/ are retained for review but never applied by the normal script.

Evidence Boundaries

  • Qwen3.5-35B physical results apply only to the exact 17,821,719,168-byte UD-IQ4_NL GGUF with SHA-256 6aec31ed654fbf12f9950c98bab0a09faea433ffdc4b3acfeb55a1f050e3e444.
  • The 8 GiB cgroup is a controlled proxy; the WSL host retained its normal physical memory.
  • Whole-device GPU memory was observed, not capped.
  • The benchmark did not capture usable cgroup io.stat bytes, so it does not establish physical SSD traffic.
  • Demand mmap is an oversubscribed long-decode mode, not a default for systems with sufficient usable RAM.
  • The retained local llama.cpp binary was built for SM89 and cannot be reused on SM120 hardware such as RTX 5090-class GPUs.

Repository Map

  • engine/ — pinned upstream identity and active patch lock
  • patches/llama.cpp/ — active, experimental, and rejected diffs
  • src/moe_cache/ — public trace, policy, and oversubscription helpers
  • scripts/ — patch, launch, provenance, benchmark, and release checks
  • tests/ — no-model correctness and release gates
  • evidence/ — sanitized summaries, never model weights or raw output
  • docs/ — reproduction, publication scope, and Chinese overview

The first public commit is a curated surface, not a dump of the private workspace. See docs/publication-scope.md.

Current Direction

llama.cpp remains the product-oriented downstream engine and the fair baseline. The project contributes measurement, provenance, focused patches, and negative evidence around memory-constrained MoE inference. A native executor remains a possible future research branch, not an implemented feature.

Models and Data

No model is bundled or automatically downloaded. Each performance claim is scoped to an exact checkpoint or GGUF, revision, quantization, and hash. Model weights, raw traces, generated work directories, credentials, private network locations, and machine-specific storage records must not be committed.

Project Policies

License

Project-authored code and documentation are available under the MIT License. Third-party code represented by the llama.cpp patch series remains subject to the upstream copyright and license notices described in THIRD_PARTY_NOTICES.md.

About

Research preview for reproducible MoE routing and memory oversubscription on consumer GPUs. Not a production inference engine.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages