Kimi K3 Is Here: Efficient Day-0 Support on vLLM
vLLM achieved efficient day-0 support for the trillion-parameter MoE model Kimi K3, paving the way for ultra-large model deployment through key optimizations like hybrid caching and speculative decoding.
vLLM Blog · Jul 27, 2026
A Preview of Production-Scale Kimi K3 Support on vLLM
vLLM, together with Moonshot AI, NVIDIA, and AMD, is racing to bring day-0 serving support for Kimi K3 by reinventing prefix caching for KDA attention and fusing kernels across hardware stacks.
vLLM Blog · Jul 22, 2026
🤗 Kernels: Major Updates
HuggingFace introduces a new 'kernel' repository type on the Hub, improves security with reproducible builds and trusted publishers, and expands framework support, laying the foundation for a standardized custom GPU kernel ecosystem.
Hugging Face Blog · Jul 6, 2026
The Zig project's rationale for their firm anti-AI contribution policy
The Zig project bans LLM-generated contributions because it invests in people, not code, believing AI assistance disrupts the process of cultivating trusted contributors.
Simon Willison · Apr 30, 2026
sqlite AGENTS.md
SQLite's AGENTS.md file sets clear boundaries for AI-generated code and bug reports, marking a shift from passive acceptance to active management of AI's impact in open-source communities.
Simon Willison ·
What We Learned by Reproducing 2,200 papers from ICML
The Hugging Face community used AI agents to reproduce about a third of ICML 2026 papers in 19 days, revealing the state of academic reproducibility and AI's new role in research auditing.
Hugging Face Blog ·