← Back to Home

Tag: 推测解码 (4 articles)

Up to 3.2x Faster Inference with LFM2.5-DSpark

Liquid AI releases DSpark draft models for its LFM2.5 family, achieving up to 3.2x faster inference through speculative decoding without compromising output quality.

Hugging Face Blog · Aug 21, 2026

Speculators v0.5.0: DFlash Support and Online Training

The Speculators v0.5.0 release introduces the DFlash algorithm for speculative decoding, which generates draft tokens in a single forward pass, significantly reducing inference latency, and unifies online and offline training workflows.

vLLM Blog ·