CANOPY Adapts Evidence Compression for Multimodal RAG
- •CANOPY adapts evidence compression to retrieved regions at multiple levels of detail
- •Five benchmarks over 33M items showed higher average accuracy than evaluated retrieval baselines
- •Compression cut reader-input evidence tokens by 14.2-27.7% at comparable accuracy in an unrouted Qwen3-VL-8B-Instruct setting
Researchers Hyojeong Yun, Jueun Kim and Wook-Shin Han at Pohang University of Science and Technology introduced CANOPY, a framework for compressing evidence after multimodal retrieval-augmented generation (RAG) retrieves text, tables, images or videos. The paper was published on October 1 and submitted to Hugging Face on October 6. CANOPY addresses how much of each retrieved item to retain: coarse units can include irrelevant material, while uniformly fine selections can remove context needed to interpret evidence.
CANOPY, short for Canonical Projection over Hierarchy, represents retrieved items as hierarchies of regions. A node encoder fine-tuned on gold evidence scores regions against a query, and a parent-relative refinement rule selects multiple regions at different levels of detail within each item. The system prunes nodes without making LLM calls. Because compression cannot recover evidence that was never retrieved, a critic requests targeted follow-up retrieval when it judges the accumulated evidence insufficient; newly retrieved items are compressed before being added.
Across five question-answering benchmarks and a heterogeneous corpus of 33M items, CANOPY achieved higher average answer accuracy than the evaluated retrieval baselines. Ablation results indicate that additional retrieval drives the main accuracy gains on questions requiring multiple evidence steps. In an unrouted Qwen3-VL-8B-Instruct setting, compression reduced evidence tokens provided to the reader by 14.2-27.7% compared with the same iterative pipeline without compression, while maintaining comparable answer accuracy. A separate author comment reports a 14-28% reduction for the same comparison.