Dynamic Image Tokenization for Efficient VLMs
Abstract
Vision Language Models (VLMs) exhibit impressive cross-modal understanding, but the large quantity of image tokens results in substantial computational overhead. Prior efforts to reduce image tokens have mainly relied on training-free approaches, which do not properly align the LLM with the compressed visual representations at inference time. We observe that, after finetuning, even the straightforward baseline of merging a fixed number of tokens at a time can outperform current training-free approaches. However, fixed pooling cannot adjust to the specific content of an image, leading to degraded performance on complex OCR tasks that require fine-grained understanding. In this paper, we introduce DRIP, a straightforward approach for dynamically tokenizing images that uses a lightweight MLP-based boundary predictor to close this performance gap. Our findings indicate that dynamic image tokenization can restore accuracy on several OCR benchmarks while preserving performance on other coarse-grained tasks. We will release the source code and model checkpoints once the paper is accepted.