MURMUR: Cross-Lingual and Multimodal Retrieval-Augmented Reasoning for Open Question Answering in Tamil and Yoruba
Abstract
As large language models (LLMs) with retrieval augmented generation (RAG) gain traction in multimodal knowledge base question answering (KBQA), concerns about their transfer to low resource languages (LRLs) remain unaddressed. We introduce MURMUR, a benchmark evaluating multimodal cross-lingual retrieval and reasoning in LRLs. Using the hardest examples from WebQA and MultimodalQA, we build a high-quality LRL benchmark through LLM-assisted translation, human validation, and culturally aligned rewriting that reflects native speaker phrasing (i.e., what a native speaker would naturally ask) while preserving answerability. We also present XM-RAG, a cross-lingual multimodal RAG pipeline for LRLs that reaches 38.1 answer accuracy, more than 6.3 points above the next best baseline. MURMUR exposes major performance gaps and failure modes in current systems. Notably, XM-RAG performs far below top English results (WebQA 64.4 and MultimodalQA 73.48), showing that existing methods still struggle with complex tasks in LRL settings. By releasing MURMUR and XM-RAG, we offer a resource to evaluate and address these gaps and guide progress toward equitable multimodal KBQA.