Addis PulseStudio

Liquid AI Ships a Speculative-Decoding Drafter for Its Vision-Language Model — and the End-to-End Numbers Are More Interesting Than the Headline

LFM2.5-VL-DSpark adds 280M parameters to the LFM2.5-VL-3B target and cuts decode time by up to 3.1× on Apple Silicon. The vision-encoder pass and prefill are untouched, so the wall-clock gain is smaller. A config change, not a rewrite.

4 min read798 words

What happened

Liquid AI published LFM2.5-VL-DSpark in September 2026: a 280M-parameter speculative-decoding drafter for its LFM2.5-VL-3B vision-language model, an 8.9% overhead on the target. Decode speedups land at 2.3× to 3.1× on Apple Silicon, with zero change to output quality because the target verifies every proposed token.

Context

LFM2.5-DSpark already existed for text-only targets; this extends the same speculative-decoding recipe to a vision-language model where image patches and text tokens share a representation space. The inference algorithm is unchanged from the text drafters, so the engineering surface is the same: swap in a VL target, swap in this drafter checkpoint. Day-one integrations ship for llama.cpp, MLX-VLM, and SGLang, and weights are published in both Safetensors and GGUF formats.

How it works

The drafter is an attention-only model with 4 layers, selected from ablations at 3, 4, and 5 tapped layers. At inference it captures the target's hidden states at those 4 layers and drafts a block of k candidate tokens (default block size 9; 8 recommended for some hardware configurations). Because image patches and text tokens are projected into a shared representation before the tapped layers, the drafter sees hidden-state vectors of identical dimensionality regardless of input modality, so the same inference path handles a chart image and a text prompt.

Training ran 10 epochs on a mixture of vision-language SFT data weighted toward expected serving workloads. At decode time the target model verifies every proposed token in one parallel pass, so greedy output is bit-identical to the target running alone. The critical constraint: speculative decoding accelerates only the decode stage. Vision encoding and prefill are untouched, so end-to-end gains are smaller than decode-only gains, especially on hardware where prefill dominates wall time.

Our read

The number that matters is the end-to-end one, not the decode one. A 3.13× decode speedup on an M5 Max becomes 1.56× to 2.62× wall-clock because the vision-encoder pass and prefill are exactly as slow as before. On an M3 Ultra the spread widens: 1.57× to 2.14× decode but only 1.30× to 1.77× end-to-end. That gap is Amdahl's law in practice. On compute-constrained hardware where prefill dominates, the drafter's contribution to perceived speed shrinks. A short single-shot caption sees roughly 1.3×; a long multi-turn video-description loop stretches toward 2.6×.

The H100 decode figure is stated in the source as "20.4× to 2.66×," a descending range that is almost certainly a typo. No per-task breakdown is published, so the 20.4× endpoint cannot be pinned to a specific workload. Treat that number as unconfirmed until Liquid AI clarifies.

What the blog omits is training-data provenance. The 10-epoch SFT mixture gets one sentence, no dataset names, no corpus licensing, no reproducibility detail. For a studio that fine-tunes or redistributes, close that gap before committing the checkpoint to a commercial pipeline.

What this changes

If a ComfyUI node calls a local VLM server over HTTP, adoption is a config change this week: pull the drafter in Safetensors or GGUF, build llama.cpp per PR #29339 or MLX-VLM per PR #2280, set block size to 8 in the config, restart. No custom inference code, no node-graph changes. The 280M-parameter overhead is negligible on a 3B target; VRAM budget is effectively unchanged.

If you run SGLang on a small GPU endpoint, the DSpark build (PR #40651) is available but one PR behind in community familiarity. For a two-person studio, the llama.cpp or MLX path is the safer bet. If ComfyUI invokes the model in-process rather than via a server, verify the llama.cpp or MLX backend is exposed through your custom node before expecting DSpark to work.

License

The source describes the model as "Open-weight — Download, fine-tune, and deploy without restrictions" but does not name a specific licence (Apache-2.0, MIT, or a Liquid AI custom term). Check the Hugging Face model card for the exact licence string before shipping anything commercial on it.

Key takeaways

  • LFM2.5-VL-DSpark is a 280M-parameter, 4-layer attention-only drafter (block size 8–9) that adds speculative decoding to LFM2.5-VL-3B with zero loss in output quality, because the target verifies every token.
  • Decode speedups range from 2.3× to 3.1× on M5 Max (MLX) and 1.6× to 2.1× on M3 Ultra (llama.cpp); end-to-end gains are 1.3× to 2.6× depending on hardware and task length.
  • The H100 decode figure is stated as "20.4× to 2.66×" in the source and reads as a probable typo; no per-task breakdown is published to verify the upper bound.
  • Day-one integrations exist for llama.cpp, MLX-VLM, and SGLang; no custom inference code is required for a local VLM serving stack.
  • The training-data mixture is described in one sentence with no dataset names or provenance, and no specific licence is named in the blog post.

Sources

  1. Accelerating vision-language models with LFM2.5-VL-DSpark — tier 1
aivision-language-modelsinference

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
1
cluster pair
gemma4:12b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20260925T223947Z