← Back to Home

Tag: 分离式推理 (1 articles)

vLLM Reaches 25K Total TPS/GPU on Qwen3.5

vLLM achieves 25K tokens/sec/GPU on Qwen3.5 through Blackwell-optimized GDN kernels, hybrid cache/state transfer, and race-free async scheduling — the optimization story goes beyond raw numbers.

vLLM Blog · Aug 6, 2026