QSA HiSparse for SGLang: 256K KV offload, CUDA Graph benchmarks, and patches tested on dual RTX 4090 48GB
qsa cpu-offload tensor-parallelism sparse-attention long-context fp8 gpu-inference llm-inference qwen sglang qwen3 rtx-4090 cuda-graphs sm89 256k-context hisparse rtx-4090-48gb 4090-48gb dual-rtx-4090 kv-cache-offloading
-
Updated
Sep 16, 2026 - Python