Benchmarking of Automated Assessment of K-5 Student Narratives Using Large Language Models
Abstract
Assessing narrative language is essential for evaluating early literacy development, but manual scoring of children’s narratives is labor-intensive and requires expert training. In this work, we study whether large language models (LLMs) can reliably assess narrative structure in K–5 student stories, focusing on rubric-based scoring of discourse elements such as character, setting, and problem. We present a systematic evaluation of LLM-based scoring across two expert-annotated datasets of transcribed oral narratives, comparing fine-tuning and prompt-only approaches across a range of modern LLM architectures. Using human inter-rater reliability (IRR) as a reference, we find that even smaller sized fine-tuned LLMs achieve near-human Quadratic Weighted Kappa (QWK) accuracy on in-distribution narratives, while state-of-the-art prompt-only approaches are less accurate. However, fine-tuned models do not generalize robustly and exhibit substantial degradation on out-of-sample story types. We find that prompt-only approach is competitive in low-resource settings, while fine-tuning is preferable when sufficient labeled data is available. This underscores a need for further research in low-resource automated narrative assessment.