Machine LearningDev6 min reading time

Up to 3.2x Faster Inference with LFM2.5-DSpark

Covered by 3 sources
Read full post
The LFM2.5 family from a research group now includes DSpark draft model checkpoints for three models, enabling up to 3.2x faster inference on GPUs and nearly 2.9x on-device without compromising output quality. DSpark uses speculative decoding with a lightweight draft model and verification to reduce latency, especially improving function-calling speed by 57%. The draft models are small (~300M parameters) and trained on diverse data, with open-source support for llama.cpp and SGLang from day one.

Covered by 3 sources


More in Machine Learning

Machine Learning4 min read

Arlequin AI raises €28M to build novel AI models that learn complex relationships at scale

SiliconANGLE
Machine Learning4 min read

DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model

The Next Web
Machine Learning3 min read

Nvidia and Palantir fine-tune a 30B Nemotron model for Nvidia’s supply chain. It beats a model 18 times its size.

Covered by 3 sources