Optimized SGLang runtime for Qwen3.8-27B FP8 with DFlash2 and Qwen3.8 Flash-Next NVFP4 with FR-Spec on one NVIDIA RTX PRO 6000 Blackwell 96 GB GPU (SM120): 524K context, HiCache and NIXL.
-
Updated
Oct 3, 2026 - Python
Optimized SGLang runtime for Qwen3.8-27B FP8 with DFlash2 and Qwen3.8 Flash-Next NVFP4 with FR-Spec on one NVIDIA RTX PRO 6000 Blackwell 96 GB GPU (SM120): 524K context, HiCache and NIXL.
INT8-quantized hierarchical KV cache for SGLang HiCache. Compresses KV evicted from GPU L1 into the CPU L2 tier, holding 1.78x more tokens in the same host memory. Measured on Qwen3-8B against a BF16 baseline.
Training-free HiCache acceleration for TRELLIS.2 image-to-3D in ComfyUI (~2x, near-lossless). Pairs with ComfyUI-Trellis2.
HiCache (Hermite) acceleration for Meta SAM 3D Objects.
HiCache++ (DMD) acceleration for SAM 3D Objects — lossless to i6.
HiCache++ (DMD) acceleration for Hunyuan3D-2.1 — lossless at larger intervals.
HiCache++ (DMD) acceleration for Hunyuan3D-2 mini — exactly lossless at i5.
HiCache (Hermite) carved-hybrid acceleration for TRELLIS.2-4B.
Training-free Hunyuan3D acceleration node for ComfyUI: skip DiT steps, forecast the flow-matching velocity (HiCache Hermite / HiCache++ DMD via hicache-pp). Pairs with ComfyUI-Hunyuan3DWrapper.
Training-free HiCache acceleration for TRELLIS image-to-3D in ComfyUI (~2x, near-lossless). Pairs with ComfyUI_TRELLIS.
HiCache (Hermite) acceleration for Hunyuan3D-2 mini.
HiCache (Hermite) acceleration for Fast-SAM3D.
HiCache++ (DMD) carved-hybrid acceleration for TRELLIS.2-4B.
HiCache (Hermite) acceleration for Hunyuan3D-2.1 image-to-3D.
HiCache++ (DMD) carved-hybrid acceleration for TRELLIS v1.
HiCache++ (DMD) acceleration for Fast-SAM3D.
To associate your repository with the hicache topic, visit your repo's landing page and select "manage topics."