Kimi K3 Is Here: Efficient Day-0 Support on vLLM
vLLM achieved efficient day-0 support for the trillion-parameter MoE model Kimi K3, paving the way for ultra-large model deployment through key optimizations like hybrid caching and speculative decoding.
vLLM Blog · Jul 27, 2026
TML Inkling on vLLM: Day-0 Support with Optimized Performance
vLLM provides day-0 support for TML Inkling, achieving 380 tok/s on 4 GB200 GPUs with full feature parity, 1M context, and multimodal input.
vLLM Blog · Jul 15, 2026
Welcome Inkling by Thinking Machines
Inkling, a 1T-parameter open model with native multimodal understanding and 1M context, redefines open-source AI with architecture innovations that enable efficient inference and agentic applications.
Hugging Face Blog · Jul 15, 2026
vLLM × HPC-Ops: High-Performance Attention and MoE Backends from Tencent Hunyuan
Tencent Hunyuan contributed production-grade high-performance Attention and MoE kernels to vLLM, achieving up to 2.95x attention speedup for mixed-length decode and 1.59x MoE speedup on H20 GPUs, reducing end-to-end latency by up to 24%.
vLLM Blog · Jul 6, 2026
EMO: Pretraining mixture of experts for emergent modularity
AI2 releases EMO, a new MoE model pretrained to enable emergent modularity, allowing users to selectively use just 12.5% of experts for a task while maintaining near full-model performance.
Hugging Face Blog · May 9, 2026