← Back to Home

Tag: 推理引擎 (3 articles)

Kimi K3 Is Here: Efficient Day-0 Support on vLLM

vLLM achieved efficient day-0 support for the trillion-parameter MoE model Kimi K3, paving the way for ultra-large model deployment through key optimizations like hybrid caching and speculative decoding.

vLLM Blog · Jul 27, 2026

A Preview of Production-Scale Kimi K3 Support on vLLM

vLLM, together with Moonshot AI, NVIDIA, and AMD, is racing to bring day-0 serving support for Kimi K3 by reinventing prefix caching for KDA attention and fusing kernels across hardware stacks.

vLLM Blog · Jul 22, 2026