MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
Abstract
As AIs are increasingly used in moral decision-making across cultural contexts, it is important for them to produce culturally aligned multilingual reasoning. However, models lack this ability, as most training and evaluation has focused on English. To this end, we present the first study on using multilingual rubrics grounded in expert-curated, theory-based dimensions from psychology and philosophy to guide reasoning. First, we propose a two-step prompting method, dubbed MERIT (Multilingual Ethics with Rubric-based, Instance-specific, Theory-grounded reasoning), using our theory-grounded rubrics: the model first generates situation-specific criteria and then uses them to produce a more coherent reasoning chain. Second, we further enhance rubric-following behavior through a self-distillation approach that does not rely on external supervision from more powerful models or human annotators, named MERIT-T (MERIT-Training). Third, to enable more accurate evaluation, we introduce MCLASH, a new multilingual benchmark constructed through careful cultural adaptation, unlike prior work that relies on direct translation. We evaluate MERIT and MERIT-T on MCLASH and previous moral decision-making benchmarks. MERIT improves average F1 by 2.71 points, and MERIT-T further improves this by 4.31 points over the base model, reaching up to 6.06 points in Korean for MCLASH. Furthermore, we reveal notable differences in the distribution of selected rubrics across languages and show that MERIT-T not only improves performance but also increases the use of native languages in the reasoning chain, enhancing legibility for non-English speakers.