Inference

Liquid AI releases a 280M draft model to accelerate LFM2.5-VL-3B

The DSpark companion model adds speculative decoding to Liquid AI’s vision-language stack with integrations for llama.cpp, MLX-VLM and SGLang.

Open Model Weights published2 Oct 2026
Primary sourceLiquid AI / Hugging Face
Source published24 Sep 2026

Liquid AI has released an experimental DSpark draft model for LFM2.5-VL-3B, extending speculative decoding from its text models into vision-language inference. The drafter is roughly 280 million parameters, or about 8.9% of the size of the 3B target model. It watches hidden states from selected layers of the target and proposes blocks of candidate tokens that the full model can verify, aiming to reduce the number of expensive decoding steps without changing the target model’s output distribution.

Liquid AI reports decode speedups of up to 3.13× on-device and 2.66× on an H100, with smaller but still substantial end-to-end gains in its tests. The source is careful to discuss limits: visual workloads include image processing and prefill costs that speculative decoding does not eliminate, so decode acceleration does not translate one-for-one into total application speed. The drafter also consumes additional memory even though it is small relative to the target.

The deployment angle is particularly relevant to an open-weight registry. Liquid AI says the release has day-one integration with llama.cpp, MLX-VLM and SGLang. That makes the new artifact both a model and a compatibility relationship: it is useful only in combination with a specific target model and a runtime that understands the speculative path. Open Model Weights should preserve that relationship explicitly, rather than implying that LFM2.5-VL-3B itself suddenly became smaller or that the reported speedups apply to every hardware configuration.

OMW REGISTRY WATCH

What this changes for the evidence layer

Compatibility and companion-weight candidate. The drafter is a separate artifact whose relationship to LFM2.5-VL-3B should be represented explicitly rather than folded into the base model.

PRIMARY SOURCE

Liquid AI / Hugging Face

This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.

Open source article ↗