NVIDIA Researchers Unify Retrieval and Long Context
- •NVIDIA researchers propose UNREAL, one model for corpus retrieval and long-context evidence selection
- •On a 21M-chunk Wikipedia index, HotpotQA recall rises from 49.1% to 73.2%
- •UNREAL improves long-context results and reduces computation from roughly 32K tokens onward
NVIDIA researchers introduced UNREAL, a model-native framework that uses one frozen large language model (LLM) to select evidence for both corpus retrieval and long-context inference. It derives retrieval queries from the model’s internal representations and adds fewer than 500K trainable parameters while leaving the backbone unchanged.
On a Wikipedia index containing 21M chunks, all four dense and hybrid UNREAL backbones outperformed state-of-the-art retriever-reranker systems. The best model increased recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue.
For long-context tasks, the same evidence-selection mechanism removes distractors before answer generation. At a maximum context length of 128K tokens, it raised NoLiMa accuracy from 1.0% to 24.83%; on LVEval at 256K, it raised F1 from 49.97% to 54.66%. Compared with processing the full context, UNREAL also reduces floating-point operations (FLOPs) and time to first token for contexts of roughly 32K tokens and longer, with larger gains as context grows. The paper says these results make model-internal evidence selection a shared approach for corpus retrieval and long-context inference with sparse evidence.
The paper was published on October 6 and submitted to Hugging Face Papers on October 7 by NVIDIA-affiliated authors led by Edan Kinderman. The listed authors are Edan Kinderman, Elad Hoffer, Yochai Blau, Brian Chmiel, Ron Banner, Daniel Soudry, and Boris Ginsburg.