← Back to Home

Tag: 推理加速 (6 articles)

Welcome Inkling by Thinking Machines

Inkling, a 1T-parameter open model with native multimodal understanding and 1M context, redefines open-source AI with architecture innovations that enable efficient inference and agentic applications.

Hugging Face Blog · Jul 15, 2026

Engineering TTS Inference in vLLM-Omni

TTS inference is a heterogeneous pipeline combining latency-bound and throughput-bound stages, making traditional LLM optimization strategies ineffective and requiring architecture-aware scheduling.

vLLM Blog · Jun 23, 2026

DiffusionGemma

Google open-sources DiffusionGemma, applying diffusion architecture to text generation for the first time, achieving over 500 tokens/sec and offering a new paradigm for high-throughput scenarios.

Simon Willison · Jun 11, 2026