TGIF: Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs
Abstract
Multimodal large language models (MLLMs) often generate confident but visually ungrounded responses, a phenomenon known as hallucination. While prior mitigation methods primarily operate at decoding time or rely on additional training, they typically expose the language model to a fixed, late-layer visual representation, overlooking the rich hierarchy encoded by vision transformers. We show that hallucination behavior is strongly influenced by the depth of visual features provided to the LLM, and that no single layer is optimal across queries. To address this, we propose Text-Guided Inter-layer Fusion (TGIF), a lightweight architectural module that dynamically reweights visual features across transformer layers based on the input text, without modifying the vision encoder or increasing the token budget. Experiments on hallucination, OCR, and general VQA benchmarks demonstrate that TGIF substantially improves visual grounding and hallucination robustness while preserving strong overall reasoning performance.