BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation
Abstract
Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive field-level errors. We construct a benchmark of 931 papers across four domains and three citation tiers---popular, low-citation, and recent post-cutoff---with version-aware ground truth. Three search-enabled frontier models (GPT-5, Claude Sonnet~4.6, Gemini~3 Flash) generate approximately 23{,}000 field-level observations. Overall accuracy is 83.6\%, but only 50.9\% of entries are fully correct; accuracy drops 27.7~pp from popular to recent papers, revealing heavy reliance on parametric memory even when search is available. Co-occurrence analysis identifies two failure modes: wholesale entry substitution and isolated field error. We present \texttt{clibib}, an open-source tool for deterministic BibTeX retrieval, as a mitigation mechanism. Two-stage integration raises accuracy to 91.5\% (+8.0~pp) and fully correct entries to 78.3\%, with a 0.8\% regression rate. Separating search from revision yields larger gains and lower regression than single-stage tool loops (0.8\% vs.\ 4.8\%), demonstrating that integration architecture matters independently of model/tool capability.