Qwen3.8-Flash-Next on 2× RTX 3090 + 128 GB RAM. 2,752 tok/s prefill · 104.5 tok/s decode (131K prompt); full 256K window at 2,654 tok/s prefill and up to 103 tok/s decode.
moe quantization mtp multi-gpu dual-gpu rtx3090 int4 fp8 vllm local-llm llm-inference qwen speculative-decoding rtx4090 rtx-3090 cpu-offloading qwen3-8 qwen38 qwen3-8-flash-next 256k-context
-
Updated
Sep 25, 2026 - Python