Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

SSR Speeds Reasoning in Multimodal Search Agents

SSR Speeds Reasoning in Multimodal Search Agents

HuggingFace·Thursday, October 8, 2026
  • •SSR replaces open-ended agent reasoning with selection from reusable language candidates at each turn
  • •Tests on seven multimodal search benchmarks used 2B and 4B models and reported competitive success rates
  • •SSR cut per-turn reasoning latency by over 90% and per-question inference latency by 28-54%
  • •SSR replaces open-ended agent reasoning with selection from reusable language candidates at each turn
  • •Tests on seven multimodal search benchmarks used 2B and 4B models and reported competitive success rates
  • •SSR cut per-turn reasoning latency by over 90% and per-question inference latency by 28-54%
  • •SSR replaces open-ended agent reasoning with selection from reusable language candidates at each turn
  • •Tests on seven multimodal search benchmarks used 2B and 4B models and reported competitive success rates
  • •SSR cut per-turn reasoning latency by over 90% and per-question inference latency by 28-54%
  • •SSR replaces open-ended agent reasoning with selection from reusable language candidates at each turn
  • •Tests on seven multimodal search benchmarks used 2B and 4B models and reported competitive success rates
  • •SSR cut per-turn reasoning latency by over 90% and per-question inference latency by 28-54%

Researchers from Carnegie Mellon University and other institutions introduced Selection-based Structured Reasoning (SSR), a framework designed to reduce the reasoning cost of multimodal search agents built with small models. The paper was published on October 1 and submitted to Hugging Face on October 7 by Xiaoyu Zhu. It says small models can produce lengthy free-form reasoning before each action, even when that reasoning offers little guidance for choosing the next action.

SSR turns reasoning into a choice among prewritten, reusable natural-language candidates. At each turn, a model selects a candidate according to its likelihood in the current context, without an auxiliary task head. The researchers use teacher-forced prefilling to score candidates in parallel, calculating token likelihoods within and across candidates while sharing a context KV cache (stored attention information reused during generation).

The team tested SSR on seven multimodal search benchmarks with 2B and 4B models, using multiple reinforcement learning objectives and supervised fine-tuning. SSR maintained an average success rate competitive with leading search agents of the same scale. It reduced per-turn reasoning latency by over 90% and total model inference latency per question by 28-54%, while preserving task performance, according to the paper's abstract. The authors describe the gains as applying across the training approaches they evaluated.

Researchers from Carnegie Mellon University and other institutions introduced Selection-based Structured Reasoning (SSR), a framework designed to reduce the reasoning cost of multimodal search agents built with small models. The paper was published on October 1 and submitted to Hugging Face on October 7 by Xiaoyu Zhu. It says small models can produce lengthy free-form reasoning before each action, even when that reasoning offers little guidance for choosing the next action.

SSR turns reasoning into a choice among prewritten, reusable natural-language candidates. At each turn, a model selects a candidate according to its likelihood in the current context, without an auxiliary task head. The researchers use teacher-forced prefilling to score candidates in parallel, calculating token likelihoods within and across candidates while sharing a context KV cache (stored attention information reused during generation).

The team tested SSR on seven multimodal search benchmarks with 2B and 4B models, using multiple reinforcement learning objectives and supervised fine-tuning. SSR maintained an average success rate competitive with leading search agents of the same scale. It reduced per-turn reasoning latency by over 90% and total model inference latency per question by 28-54%, while preserving task performance, according to the paper's abstract. The authors describe the gains as applying across the training approaches they evaluated.

Read original (English)·Oct 8, 2026
#selection based structured reasoning#multimodal search agents#inference latency#kv cache#reinforcement learning#supervised fine tuning#2b models#4b models