Reasoning over Non-Parametric Memory in Multilingual LLMs: A Cognitive-Based Analysis of Relational Knowledge
Abstract
We introduce Wiki-Bloom, a multilingual benchmark for evaluating how large language models reason over retrieved or provided relational knowledge across cognitive levels, veracity conditions, and languages, with the goal of separating reasoning from memorization. Wiki-Bloom contains 450 multiple-choice questions grounded in Wikidata relations and instantiated under factual, counterfactual, and neutralized conditions. We evaluate 12 Qwen 2.5 model variants across six languages, Arabic, Chinese, English, Hindi, Persian, and Spanish, and five Bloom-inspired cognitive levels. We find that knowledge representation matters more than either language or model scale: at 32B Instruct, neutralized contexts outperform counterfactual ones by 15.3 points, far exceeding cross-language variation. Bloom’s hierarchy also does not map cleanly onto model behavior: accuracy stays above 95% for Remember, Understand, and Apply, then drops sharply to 65.7% at Analyze and 41.3% at Evaluate. In addition, an unexpected Apply-over-Understand pattern emerges at 7B, becomes universal at 14B, and is strongest under counterfactual conditions, suggesting that structured relational reasoning may be easier for larger models than paraphrase-based understanding. Finally, error analysis shows that higher-level failures in factual settings are driven by memory interference: models often choose answers that are factually correct about an entity but relationally incorrect for the question, a failure mode that disappears when entity names are replaced with abstract placeholders. Overall, Wiki-Bloom shows that multilingual evaluation helps reveal not only cross-lingual generalization, but also how benchmark design and knowledge representation shape apparent reasoning performance. Across all six languages, the same qualitative patterns hold, indicating that they reflect intrinsic model behavior rather than language-specific artifacts.