Uno Speeds LLM Inference With Diffusion
- •Institute of Foundation Models paper introduces diffusion-augmented LLMs for parallel token sampling without quality loss
- •Uno reports higher throughput at every evaluated batch size and up to 3times speedups over base AR model
- •8B Uno outperforms 26B DiffusionGemma and Mercury 2 across agentic tool use, coding, and long-context benchmarks
Researchers from the Institute of Foundation Models published “Unlocking Lossless Speedups in LLMs via Discrete Diffusion” on Sep 3, and Hugging Face listed it as the #1 Paper of the day after Subham Sekhar Sahoo submitted it on Sep 8. The paper introduces diffusion-augmented LLMs, a model class designed to keep the probability distribution of an autoregressive language model while using diffusion to sample multiple tokens in parallel during inference. The authors say the method targets the slow, sequential token generation caused by next-token prediction and autoregressive structure.
The proposed setup separates model parameters into AR weights and lightweight diffusion weights. AR weights are trained with the standard next-token prediction objective, while diffusion weights are trained to generate multiple tokens simultaneously through a Diffusion Distillation phase that the authors describe as adding negligible overhead to existing LLM training pipelines. The system also introduces Ψ-Spec, a family of samplers that the paper says enables lossless acceleration and inference-time scaling at a fixed context length.
The resulting models are called Uno. According to the abstract, Uno can be trained from scratch or created by augmenting existing open-weight autoregressive LLMs. The authors contrast Uno with speculative decoding, saying their method does not require a separate draft model, and with diffusion LLMs, saying it accelerates generation without sacrificing the quality of the underlying autoregressive model.
The paper reports that Uno achieves higher throughput than leading speculative-decoding methods at every evaluated batch size and delivers up to 3times speedups over the base autoregressive model, including at the largest batch size supported by the device. The authors also say their 8B Uno model outperforms the leading open diffusion LLM, 26B DiffusionGemma, and proprietary Mercury 2 across all evaluated benchmarks in agentic tool use, coding, and long-context reasoning.
The Hugging Face page shows 42 upvotes, 44 collection actions, 5 models citing the paper, 0 datasets citing it, 0 Spaces citing it, and 1 collection including it. The authors listed include Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi, Varad Pimpalkhute, Yash Akhauri, Alexander Moreno, Mikhail Yurochkin, Zhenting Wang, Mostafa Elhoushi, Nolan Dey, Shane Bergsma, Joel Hestness, John Thickstun, Eric Xing, and Zhengzhong Liu.
A community discussion questioned whether Uno’s core framework resembled Orthrus, a paper linked as released 4 months earlier. Sahoo replied that Uno keeps the architecture and attention unchanged, while the linked paper changes the architecture by adding diffusion attention heads and uses bidirectional attention for diffusion blocks. Another commenter argued that both systems preserve a frozen Transformer backbone, add trainable diffusion parameters, reuse the AR KV cache, draft multiple tokens in parallel, and verify them with AR weights for lossless generation.