Up to 3.2x Faster Inference with LFM2.5-DSpark
Liquid AI releases DSpark draft models for its LFM2.5 family, achieving up to 3.2x faster inference through speculative decoding without compromising output quality.
Liquid AI releases DSpark draft models for its LFM2.5 family, achieving up to 3.2x faster inference through speculative decoding without compromising output quality.
vLLM open-sources support for parallel speculative decoding algorithms like P-EAGLE, breaking the autoregressive drafting bottleneck for higher acceptance rates and simpler tuning.
EAGLE-3 speculative decoding, deployed on AMD GPUs via vLLM, losslessly accelerates inference for large models like Kimi-K2.5, highlighting a shift toward algorithm-hardware co-design for efficient AI serving.
The Speculators v0.5.0 release introduces the DFlash algorithm for speculative decoding, which generates draft tokens in a single forward pass, significantly reducing inference latency, and unifies online and offline training workflows.