Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

V-CoLA Compresses Vision Tokens for Linear Attention

V-CoLA Compresses Vision Tokens for Linear Attention

HuggingFace·Friday, October 9, 2026
  • •V-CoLA compresses vision tokens for vision-language models using linear attention
  • •The framework retained 99.5% performance with 50.0% of vision tokens and over 88.0% with 12.5%
  • •Experiments reported prefill speedups ranging from 1.86× to 6.15× across multiple benchmarks
  • •V-CoLA compresses vision tokens for vision-language models using linear attention
  • •The framework retained 99.5% performance with 50.0% of vision tokens and over 88.0% with 12.5%
  • •Experiments reported prefill speedups ranging from 1.86× to 6.15× across multiple benchmarks
  • •V-CoLA compresses vision tokens for vision-language models using linear attention
  • •The framework retained 99.5% performance with 50.0% of vision tokens and over 88.0% with 12.5%
  • •Experiments reported prefill speedups ranging from 1.86× to 6.15× across multiple benchmarks
  • •V-CoLA compresses vision tokens for vision-language models using linear attention
  • •The framework retained 99.5% performance with 50.0% of vision tokens and over 88.0% with 12.5%
  • •Experiments reported prefill speedups ranging from 1.86× to 6.15× across multiple benchmarks

Hao Jiang and 11 coauthors proposed V-CoLA, a training-free framework for compressing vision tokens in vision-language models (VLMs), in a paper published on October 8 and submitted to Hugging Face on October 9. The method targets hybrid models with linear attention, such as Qwen3.5, where earlier compression methods designed for softmax attention may not generalize. The authors report that both attention-based and similarity-based approaches lose notable performance in this setting.

V-CoLA uses a uniqueness-aware importance criterion to identify critical vision tokens, then adaptively merges tokens to reduce their number. Its components are optimized to work with linear attention’s chunk-wise parallelism, which processes sequences in sections. The framework requires no additional training.

Across multiple benchmarks, V-CoLA retained 99.5% of original performance using 50.0% of the vision tokens. With only 12.5% of the tokens, it retained over 88.0% of original performance. The authors also report prefill speedups ranging from 1.86× to 6.15×; prefill is the stage that processes an input before a model generates its response. The paper attributes these results to experiments across multiple benchmarks but does not name the benchmarks in the supplied abstract text. The submission lists Hao Jiang, Yiru Mao, Tianpeng Bu, Hao Zhou, Hongtao Duan, Wang Jing, Bowen Xu, Xin Chen, Lulu Hu, Bin Yang, Yongliang Tao and Minying Zhang as authors.

Hao Jiang and 11 coauthors proposed V-CoLA, a training-free framework for compressing vision tokens in vision-language models (VLMs), in a paper published on October 8 and submitted to Hugging Face on October 9. The method targets hybrid models with linear attention, such as Qwen3.5, where earlier compression methods designed for softmax attention may not generalize. The authors report that both attention-based and similarity-based approaches lose notable performance in this setting.

V-CoLA uses a uniqueness-aware importance criterion to identify critical vision tokens, then adaptively merges tokens to reduce their number. Its components are optimized to work with linear attention’s chunk-wise parallelism, which processes sequences in sections. The framework requires no additional training.

Across multiple benchmarks, V-CoLA retained 99.5% of original performance using 50.0% of the vision tokens. With only 12.5% of the tokens, it retained over 88.0% of original performance. The authors also report prefill speedups ranging from 1.86× to 6.15×; prefill is the stage that processes an input before a model generates its response. The paper attributes these results to experiments across multiple benchmarks but does not name the benchmarks in the supplied abstract text. The submission lists Hao Jiang, Yiru Mao, Tianpeng Bu, Hao Zhou, Hongtao Duan, Wang Jing, Bowen Xu, Xin Chen, Lulu Hu, Bin Yang, Yongliang Tao and Minying Zhang as authors.

Read original (English)·Oct 9, 2026
#v cola#vision token compression#linear attention#vision language models#token merging#prefill speedup#qwen3.5#training free