Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

NVIDIA Researchers Unify Retrieval and Long Context

NVIDIA Researchers Unify Retrieval and Long Context

HuggingFace·Thursday, October 8, 2026
  • •NVIDIA researchers propose UNREAL, one model for corpus retrieval and long-context evidence selection
  • •On a 21M-chunk Wikipedia index, HotpotQA recall rises from 49.1% to 73.2%
  • •UNREAL improves long-context results and reduces computation from roughly 32K tokens onward
  • •NVIDIA researchers propose UNREAL, one model for corpus retrieval and long-context evidence selection
  • •On a 21M-chunk Wikipedia index, HotpotQA recall rises from 49.1% to 73.2%
  • •UNREAL improves long-context results and reduces computation from roughly 32K tokens onward
  • •NVIDIA researchers propose UNREAL, one model for corpus retrieval and long-context evidence selection
  • •On a 21M-chunk Wikipedia index, HotpotQA recall rises from 49.1% to 73.2%
  • •UNREAL improves long-context results and reduces computation from roughly 32K tokens onward
  • •NVIDIA researchers propose UNREAL, one model for corpus retrieval and long-context evidence selection
  • •On a 21M-chunk Wikipedia index, HotpotQA recall rises from 49.1% to 73.2%
  • •UNREAL improves long-context results and reduces computation from roughly 32K tokens onward

NVIDIA researchers introduced UNREAL, a model-native framework that uses one frozen large language model (LLM) to select evidence for both corpus retrieval and long-context inference. It derives retrieval queries from the model’s internal representations and adds fewer than 500K trainable parameters while leaving the backbone unchanged.

On a Wikipedia index containing 21M chunks, all four dense and hybrid UNREAL backbones outperformed state-of-the-art retriever-reranker systems. The best model increased recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue.

For long-context tasks, the same evidence-selection mechanism removes distractors before answer generation. At a maximum context length of 128K tokens, it raised NoLiMa accuracy from 1.0% to 24.83%; on LVEval at 256K, it raised F1 from 49.97% to 54.66%. Compared with processing the full context, UNREAL also reduces floating-point operations (FLOPs) and time to first token for contexts of roughly 32K tokens and longer, with larger gains as context grows. The paper says these results make model-internal evidence selection a shared approach for corpus retrieval and long-context inference with sparse evidence.

The paper was published on October 6 and submitted to Hugging Face Papers on October 7 by NVIDIA-affiliated authors led by Edan Kinderman. The listed authors are Edan Kinderman, Elad Hoffer, Yochai Blau, Brian Chmiel, Ron Banner, Daniel Soudry, and Boris Ginsburg.

NVIDIA researchers introduced UNREAL, a model-native framework that uses one frozen large language model (LLM) to select evidence for both corpus retrieval and long-context inference. It derives retrieval queries from the model’s internal representations and adds fewer than 500K trainable parameters while leaving the backbone unchanged.

On a Wikipedia index containing 21M chunks, all four dense and hybrid UNREAL backbones outperformed state-of-the-art retriever-reranker systems. The best model increased recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue.

For long-context tasks, the same evidence-selection mechanism removes distractors before answer generation. At a maximum context length of 128K tokens, it raised NoLiMa accuracy from 1.0% to 24.83%; on LVEval at 256K, it raised F1 from 49.97% to 54.66%. Compared with processing the full context, UNREAL also reduces floating-point operations (FLOPs) and time to first token for contexts of roughly 32K tokens and longer, with larger gains as context grows. The paper says these results make model-internal evidence selection a shared approach for corpus retrieval and long-context inference with sparse evidence.

The paper was published on October 6 and submitted to Hugging Face Papers on October 7 by NVIDIA-affiliated authors led by Edan Kinderman. The listed authors are Edan Kinderman, Elad Hoffer, Yochai Blau, Brian Chmiel, Ron Banner, Daniel Soudry, and Boris Ginsburg.

Read original (English)·Oct 8, 2026
#unreal#nvidia#retrieval augmented generation#long context inference#evidence selection#wikipedia index#hotpotqa#retriever reranker