Tokenisers Matter: Evaluating Bias in Multilingual and Persian-Specific Language Models
Abstract
Bias evaluation in large language models (LLMs) remains heavily English-centred, even though tokenisation, morphology, script and cultural context can substantially change what likelihood-based bias metrics measure in low-resource languages. This paper presents a tokenisation-aware evaluation of social bias in Persian, comparing LLaMA~3.1--8B, a general multilingual model, with PersianLLaMA--13B, a Persian-specific model. We construct a culturally adapted Persian StereoSet-style benchmark and a contextual dataset of 240 prompts with 720 stereotype, anti-stereotype and unrelated completions across gender, religion, race/ethnicity and profession. We evaluate models using likelihood-based scoring, Categorical Bias and an Adjusted Perplexity Index, together with explicit tokenisation diagnostics. The results show that LLaMA fragments Persian text substantially more than PersianLLaMA, which inflates variance in likelihood-based scores and can make degraded language modelling appear as either bias or neutrality. PersianLLaMA produces more compact Persian segmentation, lower directional instability and more interpretable bias profiles, especially in culturally salient categories. The findings argue that low-resource bias evaluation must jointly consider benchmark design, language-specific tokenisation and metric calibration.