Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
Abstract
Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate how reliable existing automatic MT metrics—developed for modern languages—are in this setting, using Classical Chinese–English translation as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, but MetricX is overall more sensitive to the errors that matter most. Our findings expose key limitations of current evaluation practices and motivate more robust and interpretable metrics for historically and culturally distinct translation settings.