Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents
Abstract
We study multilingual evaluation for tool-using language models at the level of action policy rather than final answer alone. Under a fixed agentic scaffold with a shared tool vocabulary, we ask whether semantically aligned tasks in different languages induce invariant action traces. We evaluate five models on nine pairwise-eligible adapted multilingual benchmarks and a controlled synthetic benchmark of 100 canonical agentic tasks translated into 23 languages. The results show only moderate cross-lingual policy stability: mean invariance is 0.591 on the adapted benchmarks and 0.578 on the synthetic benchmark. Endpoint behavior is materially less stable, with output similarity of 0.250 and 0.159, and same-answer-different-policy cases persist in both settings (2.9\% and 2.3\%). Model rankings are source-dependent, with Gemma-3-27B leading on adapted benchmarks and GPT-OSS-120B on the synthetic benchmark, and policy drift varies strongly by dataset and task family. These findings support a view of multilingual tool-using LLMs as language-conditioned policy systems rather than language-invariant reasoners. We argue that multilingual agent evaluation should treat action traces as first-class behavioral objects, not merely byproducts of prompting.