CoVLA: Vision-Language Alignment with Fewer Vision Tokens
Abstract
Vision-Language Models project every visual patch into the language model, producing hundreds of tokens that dominate sequence length despite substantial redundancy among patches. Reducing this overhead without sacrificing accuracy is challenging: existing approaches either require learned selection modules that add complexity or discard spatial structure that certain tasks depend on. We introduce CoVLA, a VLM built around the RawPool connector, which ranks patches by importance signals already present in a frozen vision encoder and forwards only the top-K patches to the language model alongside a compact global summary. Because no component of the selection mechanism is learned, the budget K can be freely changed at inference time. We further propose a curriculum training strategy that samples K uniformly during training, so that a single checkpoint exposes a controllable efficiency frontier spanning a range of token budgets. Experiments across multiple LLM scales show that CoVLA can match or exceed MLP in some settings while using substantially fewer vision tokens, with the trade-off varying by benchmark and backbone. The trade-off is task-dependent: holistic understanding benchmarks tolerate aggressive reduction, while detail-dense tasks favor full patch coverage. These findings, together with negative results on learned saliency, suggest that frozen encoder signals are a strong and sufficient basis for vision token selection.