Skip to content

Export DeepSeek V4 ratio-4 compressed sparse attention - #700

Draft
justinchuby wants to merge 1 commit into
mainfrom
deckard/deepseek-v4-csa-ratio4
Draft

Export DeepSeek V4 ratio-4 compressed sparse attention#700
justinchuby wants to merge 1 commit into
mainfrom
deckard/deepseek-v4-csa-ratio4

Conversation

@justinchuby

Copy link
Copy Markdown
Member

Summary

  • emit the frozen-v1 pkg.nxrt::CompressedSparseAttention path for DeepSeek-V4-Flash ratio-4 layers using live query-dependent indexer and compressor projections
  • emit canonical planar block-FP8 BlockQuantizedMatMul and planar-FP4 sparse BlockQuantizedMoE nodes with strict native safetensors streaming
  • thread explicit compressed/index graph state while preserving ratio-128 HCA and all non-native/non-CSA paths

Dependency and scope

Validation

  • 2135 passed, 1 skipped focused config/mapping/model/registry/GGUF/export suite
  • Ruff format check and compatible project lint rules passed; targeted mypy passed for 8 producer modules
  • real checkpoint header preflight: 46 shards, 69,187 tensors, 159,617,149,040 bytes
  • real target graph: 41 CSA ops (21 ratio-4, 20 ratio-128), 666 block-quant matmuls, 43 block-quant MoEs; 1,822 streamed targets
  • real MTP graph: 0 CSA ops, 17 block-quant matmuls, 1 block-quant MoE; 45 streamed targets

Add the ratio-4 CSA producer, canonical planar block-quant operators, strict native checkpoint streaming, and focused fail-closed coverage while preserving ratio-128 and non-CSA paths.

Co-authored-by: Copilot <[email protected]>
@github-actions

Copy link
Copy Markdown

Performance Comparison

Comparing dcf8bab0cd078d

Model Metric Baseline Current Delta
bert (feature-extraction) model_size_bytes 359 KB 359 KB +0.0%
bert (feature-extraction) num_nodes 68 68 +0.0%
falcon model_size_bytes 364 KB 364 KB +0.0%
falcon num_nodes 66 66 +0.0%
gemma2 model_size_bytes 428 KB 428 KB +0.0%
gemma2 num_nodes 105 105 +0.0%
gpt2 model_size_bytes 324 KB 324 KB +0.0%
gpt2 num_nodes 54 54 +0.0%
llama model_size_bytes 425 KB 425 KB +0.0%
llama num_nodes 60 60 +0.0%
llama (static-cache) model_size_bytes 425 KB 425 KB +0.0%
llama (static-cache) num_nodes 56 56 +0.0%
mamba (ssm-text-generation) model_size_bytes 296 KB 296 KB +0.0%
mamba (ssm-text-generation) num_nodes 94 94 +0.0%
phi3 model_size_bytes 421 KB 421 KB +0.0%
phi3 num_nodes 58 58 +0.0%
phi3 (static-cache) model_size_bytes 421 KB 421 KB +0.0%
phi3 (static-cache) num_nodes 54 54 +0.0%
qwen2 model_size_bytes 425 KB 425 KB +0.0%
qwen2 num_nodes 60 60 +0.0%
qwen2 (static-cache) model_size_bytes 425 KB 425 KB +0.0%
qwen2 (static-cache) num_nodes 56 56 +0.0%
qwen3_5_moe (hybrid-text-generation) model_size_bytes 506 KB 506 KB +0.0%
qwen3_5_moe (hybrid-text-generation) num_nodes 265 265 +0.0%
qwen3_5_text (hybrid-text-generation) model_size_bytes 458 KB 458 KB +0.0%
qwen3_5_text (hybrid-text-generation) num_nodes 127 127 +0.0%
qwen3_5_vl (hybrid-qwen-vl) model_size_bytes 977 KB 977 KB +0.0%
qwen3_5_vl (hybrid-qwen-vl) num_nodes 450 450 +0.0%
t5 (seq2seq) model_size_bytes 836 KB 836 KB +0.0%
t5 (seq2seq) num_nodes 176 176 +0.0%
whisper (speech-to-text) model_size_bytes 1008 KB 1008 KB +0.0%
whisper (speech-to-text) num_nodes 128 128 +0.0%

No performance regressions.

@github-actions

Copy link
Copy Markdown

🏗️ Architecture Diff

Comparing dcf8bab0cd078d

Model Sub-model Changes Status
bert (feature-extraction) model 0
falcon model 0
gemma2 model 0
gemma4 (gemma4) decoder 0
gemma4 (gemma4) embedding 0
gemma4 (gemma4) vision_encoder 0
gemma4_text model 0
gpt2 model 0
llama model 0
llama (static-cache) model 0
mamba (ssm-text-generation) model 0
phi3 model 0
phi3 (static-cache) model 0
qwen model 0
qwen (static-cache) model 0
qwen2 model 0
qwen2 (static-cache) model 0
qwen2_moe model 0
qwen2_moe (static-cache) model 0
qwen3 model 0
qwen3 (static-cache) model 0
qwen3_5_moe (hybrid-text-generation) model 0
qwen3_5_text (hybrid-text-generation) model 0
qwen3_5_vl (hybrid-qwen-vl) decoder 0
qwen3_5_vl (hybrid-qwen-vl) embedding 0
qwen3_5_vl (hybrid-qwen-vl) vision_encoder 0
qwen3_moe model 0
qwen3_moe (static-cache) model 0
qwen3_next (hybrid-text-generation) model 0
t5 (seq2seq) decoder 0
t5 (seq2seq) encoder 0
whisper (speech-to-text) decoder 0
whisper (speech-to-text) encoder 0

No architecture changes detected.


Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant