All-Weather VLM: Enhancing Vision Language Models' Robustness Under Adverse Imaging Conditions
Tianfu Wang ⋅ Mingyang Xie ⋅ Haoming Cai ⋅ Tianyi Xiong ⋅ Xiyao Wang ⋅ Dongdong Fu ⋅ Guan-Ming Su ⋅ Paola Cascante-Bonilla ⋅ Christopher Metzler
Abstract
Visual Language Models (VLMs) have shown strong multimodal inference capabilities, yet their robustness under real-world degradations remains underexplored. We study VLM performance under common adverse imaging conditions, including rain, fog, snow, haze, motion blur, defocus, low light, turbulence, and adherent raindrops, without prior knowledge of degradation type. Many of these conditions are critical for autonomous driving safety and related applications. We show that naive supervised fine-tuning fails to generalize across degradation types and that existing VLMs suffer substantial performance drops. A natural solution is to pre-process degraded images with restoration models before VLM inference; however, we show that these methods often hallucinate details that mislead the VLM and further hurt performance. Surprisingly, we found that skipping restoration and directly applying VLM to degraded images sometimes yields better results. Building on this finding, we propose a multi-stage training framework that integrates degradation awareness into the VLM reasoning process. A degradation-aware encoder predicts corruption type and severity to condition generation through a unified reasoning trace before the final response. Direct Preference Optimization (DPO) further improves factual grounding by treating responses from clean images as a more reliable reference for the scene's content, while penalizing degradation-induced errors. In visual question answering (VQA) tasks, our approach yields $\sim 10$ percent gains on real-world driving scenes and $\sim 4$ percent gains on real-world camera-degradation benchmarks for generic scenes, while showing that robustness can be maintained without sacrificing general capability.
Successful Page Load