LiquidAI Releases LFM2.5 DSpark Checkpoints
- •LiquidAI releases DSpark draft checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B
- •DSpark improves throughput up to 3.18x on H100 and up to 2.87x on M4 Max
- •LFM2.5-2.6B function-calling latency falls 57% on average across multi-tool scenarios
LiquidAI released DSpark draft model checkpoints on August 20, 2026 for three LFM2.5 family models: LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. The checkpoints add speculative decoding (drafting candidate tokens for verification) that LiquidAI says improves decoding speed without changing output quality, with up to 3.18x throughput improvement on a GPU, up to 2.87x on-device, and 57% average function-calling latency reduction for LFM2.5-2.6B.
DSpark targets the decode phase of LLM inference, where the article says latency usually comes from streaming weights from DRAM into SRAM rather than heavy computation. The method uses a lightweight draft model to propose candidate tokens, then has the target model verify them in one forward pass. LiquidAI says DSpark combines a DFlash-style parallel backbone, a lightweight sequential head modeled as a Markov chain (probabilistic next-state model), and a confidence-scheduled verifier that prunes low-confidence suffixes when verification would cost more than it saves.
LiquidAI trained the first DSpark draft models with a larger and more diverse data mix covering SFT, chat, code, and function-calling data. The initial drafts are simplified attention-only models with 5 layers and a block of 9; each draft model ran 15 epochs on the full dataset, and LiquidAI selected the epoch with the highest acceptance rate rather than the lowest loss. The draft models are about ~300M parameters: LFM2.5-1.2B-Instruct totals 295.7M, while LFM2.5-8B-A1B and LFM2.5-2.6B each total 327.7M. Their shared decoder stack is 241.2M, hidden-state projection is 21.0M, Markov head is 33.6M for LFM2.5-1.2B-Instruct and 65.5M for the other two, and norms plus confidence head are 27.5k.
LiquidAI says greedy decoding output remains identical to baseline greedy output because a draft token is accepted only if it matches the target model’s distribution; when rejected, the target model’s own token replaces it. The article says benchmark accuracy, including pass@1 or exact match, is unchanged by construction. Day-one support is available for llama.cpp and SGLang through upstream open-source integrations.
LiquidAI measured on-device throughput with llama.cpp and Metal on an M4 Max MacBook Pro using FP16 GGUF weights and up to 256 output tokens. GPU throughput used SGLang on a single H100 80 GB in BF16. Both test setups used DSpark block size 9, batch size 1, temperature 0, and five benchmark datasets. For LFM2.5-2.6B, mean acceptance was 4.81 of 10, H100 speedup was 2.67x from 323 to 864 tok/s, and M4 Max speedup was 2.27x from 61 to 139 tok/s. Dataset results were MATH500 5.42 acceptance, 3.06x H100 from 326 to 1000 tok/s, 2.25x M4 Max from 61 to 137 tok/s; HumanEval 4.54, 2.56x from 326 to 835 tok/s, 2.63x from 61 to 161 tok/s; MBPP 4.71, 2.64x from 326 to 861 tok/s, 2.11x from 62 to 132 tok/s; GSM8K 4.32, 2.22x from 312 to 693 tok/s, 2.36x from 60 to 143 tok/s; MT-Bench 5.07, 2.87x from 325 to 933 tok/s, 1.99x from 62 to 123 tok/s.
For LFM2.5-1.2B-Instruct, LiquidAI reported more variance in dataset acceptance rates, with speedup varying by as much as 52% depending on text distribution. Mean acceptance was 5.02 of 10, mean H100 speedup was 2.10x from 656 to 1384 tok/s, and mean M4 Max speedup was 2.54x from 138 to 350 tok/s. For LFM2.5-8B-A1B, mean acceptance rose to 6.95 of 10, mean H100 speedup was 2.54x from 418 to 1074 tok/s, and mean M4 Max speedup was 1.18x from 90 to 106 tok/s. LiquidAI attributed the smaller 18% on-device average improvement for LFM2.5-8B-A1B to the current MoE implementation in llama.cpp’s Metal backend and extra expert activation during verification.
Developers can run the DSpark draft models with SGLang using an SGLang build with DSpark support for LFM2 targets, PR #31041, and query an OpenAI-compatible endpoint at http://localhost:30000/v1. llama.cpp support requires the respective build, PR#27383. LiquidAI made the checkpoints available on Hugging Face in Safetensors and GGUF formats for LFM2.5-2.6B-DSpark, LFM2.5-1.2B-Instruct-DSpark, and LFM2.5-8B-A1B-DSpark.