Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 3 additions & 4 deletions docs/api/build_from_gguf.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ Support is capability-specific: graph import does not imply runtime packaging.

| Census | Total | Closure |
|---|---:|---|
| Architectures | 148 | graph verdicts: {'deferred': 46, 'rejected': 2, 'supported': 100}; importable: 99; quantized import: {'rejected': 41, 'supported': 107}; runtime: {'deferred': 142, 'rejected': 2, 'supported': 4} |
| Architectures | 148 | graph verdicts: {'deferred': 46, 'rejected': 2, 'supported': 100}; importable: 100; quantized import: {'rejected': 41, 'supported': 107}; runtime: {'deferred': 142, 'rejected': 2, 'supported': 4} |
| Active stored qtypes | 25 | 24 have an import route; 1 are explicitly deferred with no route |
| Serialized projector strings | 60 | {'graph-importable': 9, 'runtime-supported': 0} |
| Tokenizer pre identifiers | 87 | 56 semantic groups; route dispositions: {'deferred-compiled-semantics': 45, 'deferred-pinned-artifact-evidence': 11, 'validated-pinned-source': 31} |
Expand Down Expand Up @@ -98,13 +98,12 @@ remain machine-readable in `_route_census.py`; this table groups only shared nex
| `dependency-or-runtime-abi-blocked` | `mtp-specialized-abi` | `mtp:bailingmoe3`, `mtp:cohere2moe`, `mtp:deepseek2`, `mtp:deepseek32`, `mtp:deepseek4`, `mtp:glm-dsa`, `mtp:mimo2`, `mtp:nemotron_h_moe`, `mtp:qwen35moe`, `mtp:qwen3next`, `mtp:step35` | specialized sidecar graph; routed/cache state ABI |
| `dependency-or-runtime-abi-blocked` | `projector-runtime-abi` | `projector:resampler` | dynamic processor-to-graph media shape ABI |
| `dependency-or-runtime-abi-blocked` | `tokenizer-compiled-semantics` | `tokenizer:afmoe`, `tokenizer:bloom`, `tokenizer:chameleon`, `tokenizer:codeshell`, `tokenizer:command-r`, `tokenizer:dbrx`, `tokenizer:deepseek-coder`, `tokenizer:deepseek-llm`, `tokenizer:deepseek-v3`, `tokenizer:default`, `tokenizer:exaone`, `tokenizer:exaone-moe`, `tokenizer:falcon`, `tokenizer:gpt3-finnish`, `tokenizer:granite-docling`, `tokenizer:granite-embed-multi-97m`, `tokenizer:grok-2`, `tokenizer:hunyuan`, `tokenizer:hunyuan-dense`, `tokenizer:jais`, `tokenizer:jais-2`, `tokenizer:joyai-llm`, `tokenizer:kimi-k2`, `tokenizer:laguna`, `tokenizer:megrez`, `tokenizer:mellum2`, `tokenizer:minerva-7b`, `tokenizer:minicpm5`, `tokenizer:minimax-m2`, `tokenizer:mpt`, `tokenizer:olmo`, `tokenizer:poro-chat`, `tokenizer:refact`, `tokenizer:sarvam-moe`, `tokenizer:seed-coder`, `tokenizer:smaug-bpe`, `tokenizer:solar-open`, `tokenizer:stablelm2`, `tokenizer:starcoder`, `tokenizer:superbpe`, `tokenizer:tekken`, `tokenizer:trillion`, `tokenizer:viking`, `tokenizer:whitespace`, `tokenizer:youtu` | compiled pinned llama.cpp oracle; dispatch-equivalence fixture |
| `evidence-only` | `architecture-runtime-evidence` | `architecture:apertus`, `architecture:arcee`, `architecture:arctic`, `architecture:baichuan`, `architecture:bailingmoe`, `architecture:bert`, `architecture:bitnet`, `architecture:bloom`, `architecture:chatglm`, `architecture:codeshell`, `architecture:cohere2`, `architecture:command-r`, `architecture:dbrx`, `architecture:deci`, `architecture:deepseek`, `architecture:dflash`, `architecture:dots1`, `architecture:dream`, `architecture:eagle3`, `architecture:ernie4_5`, `architecture:ernie4_5-moe`, `architecture:eurobert`, `architecture:exaone`, `architecture:falcon`, `architecture:gemma`, `architecture:gemma-embedding`, `architecture:gemma2`, `architecture:gemma3`, `architecture:gemma4`, `architecture:gpt2`, `architecture:gptneox`, `architecture:granite`, `architecture:granitemoe`, `architecture:hunyuan-dense`, `architecture:hy_v3`, `architecture:internlm2`, `architecture:jais`, `architecture:jais2`, `architecture:jina-bert-v2`, `architecture:jina-bert-v3`, `architecture:lfm2moe`, `architecture:llada`, `architecture:llada-moe`, `architecture:llama-embed`, `architecture:maincoder`, `architecture:mamba`, `architecture:mamba2`, `architecture:minicpm`, `architecture:minicpm3`, `architecture:modern-bert`, `architecture:mpt`, `architecture:muse-glimmer`, `architecture:nemotron`, `architecture:nemotron_h`, `architecture:neo-bert`, `architecture:nomic-bert`, `architecture:nomic-bert-moe`, `architecture:olmo`, `architecture:olmo2`, `architecture:olmoe`, `architecture:openelm`, `architecture:orion`, `architecture:pangu-embedded`, `architecture:phi2`, `architecture:phi3`, `architecture:phimoe`, `architecture:plamo`, `architecture:plm`, `architecture:qwen`, `architecture:qwen2moe`, `architecture:qwen2vl`, `architecture:qwen3`, `architecture:qwen35`, `architecture:qwen3moe`, `architecture:qwen3next`, `architecture:refact`, `architecture:rnd1`, `architecture:seed_oss`, `architecture:smallthinker`, `architecture:smollm3`, `architecture:stablelm`, `architecture:starcoder`, `architecture:starcoder2`, `architecture:t5`, `architecture:t5encoder`, `architecture:talkie`, `architecture:xverse` | immutable representative GGUF; full-logit prefill and cached-decode parity; deterministic generation/state evidence |
| `evidence-only` | `architecture-runtime-evidence` | `architecture:apertus`, `architecture:arcee`, `architecture:arctic`, `architecture:baichuan`, `architecture:bailingmoe`, `architecture:bert`, `architecture:bitnet`, `architecture:bloom`, `architecture:chatglm`, `architecture:codeshell`, `architecture:cohere2`, `architecture:command-r`, `architecture:dbrx`, `architecture:deci`, `architecture:deepseek`, `architecture:dflash`, `architecture:dots1`, `architecture:dream`, `architecture:eagle3`, `architecture:ernie4_5`, `architecture:ernie4_5-moe`, `architecture:eurobert`, `architecture:exaone`, `architecture:falcon`, `architecture:gemma`, `architecture:gemma-embedding`, `architecture:gemma2`, `architecture:gemma3`, `architecture:gemma4`, `architecture:glm-dsa`, `architecture:gpt2`, `architecture:gptneox`, `architecture:granite`, `architecture:granitemoe`, `architecture:hunyuan-dense`, `architecture:hy_v3`, `architecture:internlm2`, `architecture:jais`, `architecture:jais2`, `architecture:jina-bert-v2`, `architecture:jina-bert-v3`, `architecture:lfm2moe`, `architecture:llada`, `architecture:llada-moe`, `architecture:llama-embed`, `architecture:maincoder`, `architecture:mamba`, `architecture:mamba2`, `architecture:minicpm`, `architecture:minicpm3`, `architecture:modern-bert`, `architecture:mpt`, `architecture:muse-glimmer`, `architecture:nemotron`, `architecture:nemotron_h`, `architecture:neo-bert`, `architecture:nomic-bert`, `architecture:nomic-bert-moe`, `architecture:olmo`, `architecture:olmo2`, `architecture:olmoe`, `architecture:openelm`, `architecture:orion`, `architecture:pangu-embedded`, `architecture:phi2`, `architecture:phi3`, `architecture:phimoe`, `architecture:plamo`, `architecture:plm`, `architecture:qwen`, `architecture:qwen2moe`, `architecture:qwen2vl`, `architecture:qwen3`, `architecture:qwen35`, `architecture:qwen3moe`, `architecture:qwen3next`, `architecture:refact`, `architecture:rnd1`, `architecture:seed_oss`, `architecture:smallthinker`, `architecture:smollm3`, `architecture:stablelm`, `architecture:starcoder`, `architecture:starcoder2`, `architecture:t5`, `architecture:t5encoder`, `architecture:talkie`, `architecture:xverse` | immutable representative GGUF; full-logit prefill and cached-decode parity; deterministic generation/state evidence |
| `evidence-only` | `draft-runtime-evidence` | `draft:dflash`, `draft:eagle3` | target acceptance loop; draft cache orchestration; deterministic speedup parity |
| `evidence-only` | `mtp-runtime-evidence` | `mtp:hy_v3`, `mtp:qwen35` | target acceptance loop; cache-threaded draft/target parity |
| `evidence-only` | `projector-runtime-evidence` | `projector:adapter`, `projector:gemma3`, `projector:gemma4v`, `projector:ldp`, `projector:ldpv2`, `projector:mlp`, `projector:muse-glimmer`, `projector:qwen2.5vl_merger`, `projector:qwen2vl_merger` | paired text target; processor boundary; deterministic multimodal package execution |
| `evidence-only` | `tokenizer-artifact-evidence` | `tokenizer:bailingmoe`, `tokenizer:bailingmoe2`, `tokenizer:chatglm-bpe`, `tokenizer:cohere2moe`, `tokenizer:glm4`, `tokenizer:llada-moe`, `tokenizer:tiny_aya` | immutable GGUF/source pair; ordered vocabulary and encoding parity |
| `immediately-implementable` | `architecture-implementation` | `architecture:grok`, `architecture:grovemoe`, `architecture:hunyuan-moe`, `architecture:minimax-m2`, `architecture:mistral4` | exact metadata extraction; tensor closure; dedicated graph and parity |
| `immediately-implementable` | `architecture-implementation` | `architecture:glm-dsa` | suffix-exact tensor mapping; packed-value transform proof; graph closure |
| `immediately-implementable` | `projector-implementation` | `projector:cogvlm`, `projector:deepseekocr`, `projector:deepseekocr2`, `projector:dots3note_a`, `projector:dots3note_v`, `projector:dots_ocr`, `projector:exaone4_5`, `projector:gemma3na`, `projector:gemma3nv`, `projector:gemma4a`, `projector:gemma4ua`, `projector:gemma4uv`, `projector:glm4v`, `projector:glma`, `projector:granite4_vision`, `projector:granite_speech`, `projector:hunyuanvl`, `projector:idefics3`, `projector:internvl`, `projector:janus_pro`, `projector:kimik25`, `projector:kimivl`, `projector:lfm2`, `projector:lfm2a`, `projector:lightonocr`, `projector:llama4`, `projector:meralion`, `projector:mimo_audio`, `projector:mimovl`, `projector:minicpmv4_6`, `projector:minimax_m3`, `projector:musicflamingo`, `projector:nemotron_v2_vl`, `projector:paddleocr`, `projector:parakeet`, `projector:pixtral`, `projector:pockettts_spkenc`, `projector:qwen2.5o`, `projector:qwen2a`, `projector:qwen3a`, `projector:qwen3tts_spkenc`, `projector:qwen3vl_merger`, `projector:step3vl`, `projector:ultravox`, `projector:voxtral`, `projector:yasa2`, `projector:youtuvl` | metadata schema; tensor closure; component graph parity |
| `intentionally-rejected` | `policy-rejections` | `architecture:bailingmoe2`, `architecture:clip`, `architecture:dots3note`, `architecture:exaone-moe`, `architecture:exaone4`, `architecture:glm4`, `architecture:glm4moe`, `architecture:gptj` | policy change plus independent correctness proof |
| `intentionally-rejected` | `policy-rejections` | `projector:pockettts_gen`, `projector:qwen3tts_gen` | sidecar role must become a valid projector contract |
Expand Down Expand Up @@ -199,7 +198,7 @@ Reason codes are concise user-facing categories; detailed architecture audits re
| `gemma3n` | — | none (fails before config extraction) | exact-direct-loader-conditional-union | config=deferred; tensor_map=deferred; graph=deferred; runtime=deferred; quantized_import=supported | CONFIG_DEFERRED — Gemma3n GGUF is the text member of a vision-and-audio package whose gemma3nv and gemma3na clip companions carry distinct encoders and projectors. |
| `gemma4` | — | model=`gemma4_text`; tensor=`llama`+`gemma4_extras`; mmproj=`gemma4` | not claimed | config=supported; tensor_map=supported; graph=supported; runtime=deferred; quantized_import=supported | RUNTIME_EVIDENCE_PENDING — Config extraction, exact tensor-name closure, and a full synthetic GGUF graph build are covered, but no representative real-weight GGUF has yet passed ORT parity or generation validation. |
| `gemma4-assistant` | — | none (fails before config extraction) | audited-direct-loader-conditional-union | config=deferred; tensor_map=deferred; graph=deferred; runtime=deferred; quantized_import=supported | CONFIG_DEFERRED — Gemma4 Assistant is a standalone target-coupled model with pre/post projections, masked embeddings, scalar layer scales, its own KV cache, and a live target-model context. |
| `glm-dsa` | `glm_dsa` | none (no tensor mapping route) | audited-direct-loader-conditional-union | config=supported; tensor_map=deferred; graph=supported; runtime=deferred; quantized_import=supported | TENSOR_MAP_DEFERRED — Config extraction and the glm_moe_dsa graph are both available, but no GGUF→HuggingFace tensor-name mapping has been written for GLM-5.2's MLA + DSA-indexer tensor families yet, so weights cannot be routed into the graph. |
| `glm-dsa` | `glm_dsa` | model=`glm_moe_dsa`; tensor=`glm_dsa` | audited-direct-loader-conditional-union | config=supported; tensor_map=supported; graph=supported; runtime=deferred; quantized_import=supported | RUNTIME_EVIDENCE_PENDING — Config extraction, exact tensor-name closure, and a full synthetic GGUF graph build are covered, but no representative real-weight GGUF has yet passed ORT parity or generation validation. |
| `glm4` | — | none (fails before config extraction) | audited-direct-loader-conditional-union | config=deferred; tensor_map=deferred; graph=deferred; runtime=deferred; quantized_import=supported | CONFIG_DEFERRED — GLM4 serializes complete fused-FFN trailing blocks and NextN tensors, but the pinned loader skips appended blocks; GLM-OCR converter transforms also permute Q/K for M-RoPE. |
| `glm4moe` | — | none (fails before config extraction) | audited-direct-loader-conditional-union | config=deferred; tensor_map=deferred; graph=deferred; runtime=deferred; quantized_import=supported | CONFIG_DEFERRED — GLM4-MoE serializes biased attention and periodic dense/routed expert trailing blocks with mandatory router bias, but the pinned loader skips them. |
| `gpt-oss` | — | none (fails before config extraction) | not claimed | config=deferred; tensor_map=deferred; graph=deferred; runtime=deferred; quantized_import=supported | CONFIG_DEFERRED — The pinned GPT-OSS converter splits interleaved gate/up expert rows and repacks checkpoint block+scale tensors into expert-major MXFP4 values. |
Expand Down
60 changes: 42 additions & 18 deletions src/mobius/components/_deepseek_mla.py
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,7 @@ def __init__(
config: ArchitectureConfig,
scale: float | None = None,
linear_class: type | None = None,
split_kv_b: bool = False,
):
super().__init__()
if linear_class is None:
Expand Down Expand Up @@ -82,12 +83,25 @@ def __init__(
bias=False,
)
self.kv_a_layernorm = RMSNorm(self.kv_lora_rank, eps=config.rms_norm_eps)
# Decompresses latent KV into per-head k_nope + v
self.kv_b_proj = linear_class(
self.kv_lora_rank,
self.num_heads * (self.qk_nope_head_dim + self.v_head_dim),
bias=False,
)
self._split_kv_b = split_kv_b
if split_kv_b:
self.k_b_proj = linear_class(
self.kv_lora_rank,
self.num_heads * self.qk_nope_head_dim,
bias=False,
)
self.v_b_proj = linear_class(
self.kv_lora_rank,
self.num_heads * self.v_head_dim,
bias=False,
)
else:
# Decompresses latent KV into per-head k_nope + v
self.kv_b_proj = linear_class(
self.kv_lora_rank,
self.num_heads * (self.qk_nope_head_dim + self.v_head_dim),
bias=False,
)

self.o_proj = linear_class(
self.num_heads * self.v_head_dim,
Expand Down Expand Up @@ -154,18 +168,28 @@ def forward(

# Decompress latent KV → per-head k_nope + v
k_pass = self.kv_a_layernorm(op, k_pass)
kv_decompressed = self.kv_b_proj(op, k_pass)
# (B, S, num_heads * (nope + v_dim)) → (B, S, num_heads, nope + v_dim)
kv_decompressed = op.Reshape(
kv_decompressed,
[0, 0, self.num_heads, self.qk_nope_head_dim + self.v_head_dim],
)
k_nope, value_states = op.Split(
kv_decompressed,
[self.qk_nope_head_dim, self.v_head_dim],
axis=-1,
_outputs=2,
)
if self._split_kv_b:
k_nope = op.Reshape(
self.k_b_proj(op, k_pass),
[0, 0, self.num_heads, self.qk_nope_head_dim],
)
value_states = op.Reshape(
self.v_b_proj(op, k_pass),
[0, 0, self.num_heads, self.v_head_dim],
)
else:
kv_decompressed = self.kv_b_proj(op, k_pass)
# (B, S, num_heads * (nope + v_dim)) → (B, S, num_heads, nope + v_dim)
kv_decompressed = op.Reshape(
kv_decompressed,
[0, 0, self.num_heads, self.qk_nope_head_dim + self.v_head_dim],
)
k_nope, value_states = op.Split(
kv_decompressed,
[self.qk_nope_head_dim, self.v_head_dim],
axis=-1,
_outputs=2,
)
# k_nope: (B, S, H, nope_dim) → (B, S, H*nope_dim)... not needed yet
# value_states: (B, S, H, v_dim) → (B, S, H*v_dim) for Attention op
value_states = op.Reshape(value_states, [0, 0, -1])
Expand Down
12 changes: 5 additions & 7 deletions src/mobius/integrations/gguf/_arch_registry.py
Original file line number Diff line number Diff line change
Expand Up @@ -2079,13 +2079,11 @@
model_type="glm_moe_dsa",
aliases=frozenset({"glm_dsa"}),
config_key_map="glm_dsa",
tensor_map=Support.DEFERRED,
reason=(
"Config extraction and the glm_moe_dsa graph are both available, but "
"no GGUF→HuggingFace tensor-name mapping has been written for GLM-5.2's "
"MLA + DSA-indexer tensor families yet, so weights cannot be routed "
"into the graph. " + _NO_TENSOR_MAP
),
config_postprocessor="glm_dsa",
tensor_map_recipe=("glm_dsa",),
tensor_processor="glm_dsa",
rope_interleave=True,
reason=_RUNTIME_VALIDATION_PENDING,
),
GGUFArchitectureSpec(
gguf_arch="apertus",
Expand Down
3 changes: 2 additions & 1 deletion src/mobius/integrations/gguf/_arch_registry_test.py
Original file line number Diff line number Diff line change
Expand Up @@ -60,7 +60,7 @@
#: Number of importable architectures. Pinned so that adding support is a
#: deliberate act that also updates the documented support matrix, and so that
#: accidentally losing an architecture is a failure rather than a silence.
_EXPECTED_SUPPORTED_COUNT = 99
_EXPECTED_SUPPORTED_COUNT = 100
_PROMOTED_CONVENTIONAL_DECODERS = frozenset(
{
"bitnet",
Expand Down Expand Up @@ -190,6 +190,7 @@
"gemma3",
"gemma4",
"granite",
"glm-dsa",
"granitemoe",
"hunyuan-dense",
"jamba",
Expand Down
Loading
Loading