When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies
Abhinav Havaldar ⋅ Enrico Santus
Abstract
Retrieval-augmented generation (RAG) is widely assumed to mitigate factual errors in large language models (LLMs), but it remains unclear whether retrieval uniformly compensates for missing knowledge. We study this question in a controlled factual QA setting over public companies, constructing a benchmark of $\sim$2,000 firms across global equity indices. We evaluate six LLMs on six atomic attributes under three conditions: no context, correct context, and misleading context. We find strong geographic disparities in no-context accuracy, indicating uneven parametric knowledge. While correct context improves performance, it does not eliminate these gaps: gains are correlated with baseline accuracy, suggesting retrieval effectiveness is coupled to internal representations. Under misleading context, models frequently copy incorrect information. Larger models improve overall performance but do not remove these structural effects. These results challenge the view of RAG as a universal corrective and highlight the interaction between model knowledge, context quality, and entity representation.
Successful Page Load