Liquid AI Ships a Speculative-Decoding Drafter for Its Vision-Language Model — and the End-to-End Numbers Are More Interesting Than the Headline
LFM2.5-VL-DSpark adds 280M parameters to the LFM2.5-VL-3B target and cuts decode time by up to 3.1× on Apple Silicon. The vision-encoder pass and prefill are untouched, so the wall-clock gain is smaller. A config change, not a rewrite.
What happened
Liquid AI published LFM2.5-VL-DSpark in September 2026: a 280M-parameter speculative-decoding drafter for its LFM2.5-VL-3B vision-language model, an 8.9% overhead on the target. Decode speedups land at 2.3× to 3.1× on Apple Silicon, with zero change to output quality because the target verifies every proposed token.
Context
LFM2.5-DSpark already existed for text-only targets; this extends the same speculative-decoding recipe to a vision-language model where image patches and text tokens share a representation space. The inference algorithm is unchanged from the text drafters, so the engineering surface is the same: swap in a VL target, swap in this drafter checkpoint. Day-one integrations ship for llama.cpp, MLX-VLM, and SGLang, and weights are published in both Safetensors and GGUF formats.
How it works
The drafter is an attention-only model with 4 layers, selected from ablations at 3, 4, and 5 tapped layers. At inference it captures the target's hidden states at those 4 layers and drafts a block of k candidate tokens (default block size 9; 8 recommended for some hardware configurations). Because image patches and text tokens are projected into a shared representation before the tapped layers, the drafter sees hidden-state vectors of identical dimensionality regardless of input modality, so the same inference path handles a chart image and a text prompt.
Training ran 10 epochs on a mixture of vision-language SFT data weighted toward expected serving workloads. At decode time the target model verifies every proposed token in one parallel pass, so greedy output is bit-identical to the target running alone. The critical constraint: speculative decoding accelerates only the decode stage. Vision encoding and prefill are untouched, so end-to-end gains are smaller than decode-only gains, especially on hardware where prefill dominates wall time.
Our read
The number that matters is the end-to-end one, not the decode one. A 3.13× decode speedup on an M5 Max becomes 1.56× to 2.62× wall-clock because the vision-encoder pass and prefill are exactly as slow as before. On an M3 Ultra the spread widens: 1.57× to 2.14× decode but only 1.30× to 1.77× end-to-end. That gap is Amdahl's law in practice. On compute-constrained hardware where prefill dominates, the drafter's contribution to perceived speed shrinks. A short single-shot caption sees roughly 1.3×; a long multi-turn video-description loop stretches toward 2.6×.
The H100 decode figure is stated in the source as "20.4× to 2.66×," a descending range that is almost certainly a typo. No per-task breakdown is published, so the 20.4× endpoint cannot be pinned to a specific workload. Treat that number as unconfirmed until Liquid AI clarifies.
What the blog omits is training-data provenance. The 10-epoch SFT mixture gets one sentence, no dataset names, no corpus licensing, no reproducibility detail. For a studio that fine-tunes or redistributes, close that gap before committing the checkpoint to a commercial pipeline.
What this changes
If a ComfyUI node calls a local VLM server over HTTP, adoption is a config change this week: pull the drafter in Safetensors or GGUF, build llama.cpp per PR #29339 or MLX-VLM per PR #2280, set block size to 8 in the config, restart. No custom inference code, no node-graph changes. The 280M-parameter overhead is negligible on a 3B target; VRAM budget is effectively unchanged.
If you run SGLang on a small GPU endpoint, the DSpark build (PR #40651) is available but one PR behind in community familiarity. For a two-person studio, the llama.cpp or MLX path is the safer bet. If ComfyUI invokes the model in-process rather than via a server, verify the llama.cpp or MLX backend is exposed through your custom node before expecting DSpark to work.
License
The source describes the model as "Open-weight — Download, fine-tune, and deploy without restrictions" but does not name a specific licence (Apache-2.0, MIT, or a Liquid AI custom term). Check the Hugging Face model card for the exact licence string before shipping anything commercial on it.
Key takeaways
- LFM2.5-VL-DSpark is a 280M-parameter, 4-layer attention-only drafter (block size 8–9) that adds speculative decoding to LFM2.5-VL-3B with zero loss in output quality, because the target verifies every token.
- Decode speedups range from 2.3× to 3.1× on M5 Max (MLX) and 1.6× to 2.1× on M3 Ultra (llama.cpp); end-to-end gains are 1.3× to 2.6× depending on hardware and task length.
- The H100 decode figure is stated as "20.4× to 2.66×" in the source and reads as a probable typo; no per-task breakdown is published to verify the upper bound.
- Day-one integrations exist for llama.cpp, MLX-VLM, and SGLang; no custom inference code is required for a local VLM serving stack.
- The training-data mixture is described in one sentence with no dataset names or provenance, and no specific licence is named in the blog post.
Sources
How this post was made
Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.
- Drafted
- Independent sources
- 1
- cluster pair
- gemma4:12b
- cluster label
- gemma4:12b
- radar brief
- gemma4:12b
- research brief
- qwen3.8:27b
- draft article
- qwen3.8:27b
- short script
- qwen3.8:27b
- seo pack
- gemma4:12b
- Run
- editorial-20260925T223947Z