Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
Abstract
Multimodal Large Language Models (MLLMs) achieve remarkable performance for fine-grained pixel-level understanding tasks. However, all SOTA methods for pixel-grounded understanding rely heavily on extra components, such as a vision encoder (CLIP), segmentation experts, a mask tokenizer, and detokenizer, leading to high system complexity and limiting model scaling. In this work, we aim to explore how to build a highly simplified MLLM without introducing extra components and achieve SOTA performance for pixel-grounded tasks. Our work is motivated by the recent works on Single trAnsformer as a unified vIsion-Language Model (SAIL) design, where these works jointly learn vision tokens and text tokens in transformers. We discover that the single transformer architecture can deeply understand visual signals, not only including the alignment and transformation between visual signals and language signals, but also including the ability of fine-grained pixel-level understanding. Based on this finding, we present Pixel-SAIL, a single transformer for pixel-wise MLLM tasks. Pixel-SAIL contains some simple but effective designs, including a visual prompt injection strategy, a learnable upsampling module, and a vision expert distillation strategy to enhance pixel-level perception and visual prompt understanding capabilities. In addition, we have collected a comprehensive pixel understanding benchmark (PerBench), using a manual check. It includes three tasks: detailed object description, visual prompt-based question answering, and visual-text referring segmentation. Extensive experiments on referring segmentation, visual prompt understanding, and our benchmarks show that Pixel-SAIL achieves comparable or even better results with a much simpler architecture.