TML Inkling on vLLM: Day-0 Support with Optimized Performance
vLLM provides day-0 support for TML Inkling, achieving 380 tok/s on 4 GB200 GPUs with full feature parity, 1M context, and multimodal input.
vLLM Blog · Jul 15, 2026
vLLM provides day-0 support for TML Inkling, achieving 380 tok/s on 4 GB200 GPUs with full feature parity, 1M context, and multimodal input.
Z.ai releases GLM-5.2, the first open-source model to achieve stable 1M-token context and rival top closed-source models on long-horizon coding benchmarks.
DeepSeek V4 achieves efficient million-token long-context inference on vLLM through innovative KV cache compression and sparse attention mechanisms, marking a new era for long-text processing.