What If Chinese Were Latinized? A Counterfactual Study of Script, Tokenization, and Language Modeling
Abstract
If Chinese had been Latinized, as the Latinxua Sin Wenz movement proposed, what would it have meant for Chinese NLP? Homophones and tone marking complicate the answer. We convert Chinese corpora to syllable-separated Pinyin and train superBPE tokenizers from scratch. The resulting Pinyin vocabularies not only exhibit expected homophone collisions, but also show higher fertility. Under matched-token compute, we further pretrain a Llama-style model on Chinese characters and on Pinyin; the Pinyin model is worse on aggregate across per-character perplexity, homophone disambiguation, and overall ZhoBLiMP minimal-pair accuracy. The broader implication is that writing systems are not neutral encodings but part of the linguistic representation for language models. This perspective extends naturally to cases such as Vietnamese tone diacritics and Korean Hangul, where script should be accounted for in cross-lingual fairness.