SearchGen Framework Improves Accuracy in Agentic Visual Generation
- •TIGER-Lab introduces SearchGen to enable agentic visual generation using external multimodal search.
- •Frontier generators score only 21 to 28 out of 100 on the new 20,839-prompt SearchGen-Bench.
- •The team released a 1M-item search corpus and co-training framework to improve world-knowledge grounding.
Visual generators often struggle with long-tailed or evolving knowledge, frequently fabricating details when prompts exceed their training data. Researchers from TIGER-Lab have introduced SearchGen, a search-augmented framework designed to enable agentic visual generation by retrieving relevant multimodal context. The study establishes SearchGen-20K, a benchmark containing 20,839 prompts across 12 failure categories and 22 domains, to test how well models can integrate external information into image generation. Evaluated against this benchmark, current frontier image generators score only 21 to 28 out of 100, representing a significant performance collapse that remains hidden in existing evaluation standards.
The research identifies a structural bottleneck in how generators handle knowledge: the fixed nature of their training corpora versus the open-ended requirements of real-world visual requests. While naive search tools can inject noise, the authors demonstrate that an agent can actively retrieve and organize information. They propose a teach-then-search co-training framework that allows models to learn the boundary between what they can internalize and what must be retrieved from external sources. Even a minimal version of this recipe produces monotonic improvements, establishing a foundation for recursive self-improvement in generation tasks.
To support further research, the team has released the SearchGen-Bench dataset, a reproducible SearchGen-Corpus-1M for offline testing, and the complete co-training dataset. These resources serve as a harness for developers building world-knowledge-grounded visual generation systems. The framework shifts the role of LLM agents from simple prompt rewriting to active orchestration of multimodal context, ensuring images align with specific, real-world informational needs that fixed training data cannot address.