面向 8×RTX 4090(SM89)的 DeepSeek-V4-Flash-0731 MXFP4 SGLang 实验分支
-
Updated
Sep 14, 2026 - Python
面向 8×RTX 4090(SM89)的 DeepSeek-V4-Flash-0731 MXFP4 SGLang 实验分支
QSA HiSparse for SGLang: 256K KV offload, CUDA Graph benchmarks, and patches tested on dual RTX 4090 48GB
Feeding the Tensor Cores: a dense FP16 GEMM for NVIDIA Ada (sm_89) at 96.5% of cuBLAS, and what it teaches about how each GPU generation handles async copy and Tensor Core issue.
Run DeepSeek-V4-Flash with 1M-token context on pre-Hopper NVIDIA GPUs (sm_89) — a vLLM 0.26.0 overlay replacing all SM90+ DeepGEMM/FlashMLA operators with Triton/TileLang/Marlin kernels + CPU-KV-pool (pinned host memory via CUDA UVA).
To associate your repository with the sm89 topic, visit your repo's landing page and select "manage topics."