Qwen3.8-Flash-Next on 2× RTX 3090 + 128 GB RAM: 1,402 tok/s prefill, 135.2 tok/s decode, and full 256K context.
moe quantization mtp multi-gpu dual-gpu rtx3090 int4 fp8 vllm local-llm llm-inference qwen speculative-decoding rtx4090 rtx-3090 cpu-offloading qwen3-8 qwen38 qwen3-8-flash-next 256k-context
-
Updated
Sep 18, 2026 - Python