We Hebben Een Serieus Translatie: Modeling Intercomprehension as Probabilistic Inference
Abstract
As a result of language evolution, some language pairs are more closely related than others, leading to varying degrees of mutual intelligibility or intercomprehension between languages: a speaker of L1 may partially understand utterances in L2 despite having no prior exposure to L2. How is this zero-shot cross-language comprehension possible? In this work, we extend past work on algorithmic models of noisy-channel inference to model mutual intelligibility in a resource-rational, Bayesian framework. The model uses an LM in L1 only for scoring latent hypotheses about the translations of observed L2 utterances, and a general-purpose error model to infer a mapping between L2 and L1 words based on either form-based similarity or memorized rules. We also conduct a human experiment, eliciting inferences from monolingual English speakers for simple Dutch sentences. Our full model shows a closer alignment to the distribution of human intercomprehension performance than ablations, and captures variance in item-level difficulty. These results provide a cognitively plausible computational model of intercomprehension, and highlight the flexible inferences made by comprehenders under wide uncertainty in real-world cross-language scenarios.