A Mirage of Coherence: How Metaphor Impacts Language Models' Discourse Coherence Assessment
Abstract
Recent approaches to discourse coherence assessment increasingly rely on language models, but it remains unclear whether they evaluate underlying discourse structure or merely surface-level lexical patterns. To investigate, we introduce metaphor as a controlled linguistic probe. Using a rigorous paraphrase framework, we generate meaning-preserving metaphorical rewrites for a corpus of human-annotated texts. Across multiple architectures, we uncover a pervasive tendency: models paradoxically reward figurative language, consistently assigning higher coherence scores to texts with novel metaphors, provided the metaphors remain semantically compatible with their context. Furthermore, layer-wise analyses reveal that while early model layers are highly sensitive to surface-level figurative features, higher layers successfully recover the underlying semantics. Ultimately, our findings demonstrate that coherence judgments are influenced by local stylistic variation, indicating that current models often evaluate linguistic realization rather than relying on discourse-level reasoning alone.