← Back to Home

Tag: 注意力机制 (5 articles)

A Preview of Production-Scale Kimi K3 Support on vLLM

vLLM, together with Moonshot AI, NVIDIA, and AMD, is racing to bring day-0 serving support for Kimi K3 by reinventing prefix caching for KDA attention and fusing kernels across hardware stacks.

vLLM Blog · Jul 22, 2026

The Transformer Family Version 2.0

Lilian Weng's new article deeply explores the evolution and new features of Transformers, revealing their ongoing impact in natural language processing.

Lilian Weng · Jan 27, 2023

Which tokens does a hybrid model predict better?

Hybrid models significantly outperform pure Transformers in semantic understanding and dynamic context tracking, but lag in verbatim repetition, revealing a clear architectural division of labor.

Hugging Face Blog ·