BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence
Sean Wu ⋅ Fredrik K. Gustafsson ⋅ Edward Phillips ⋅ Boyan Gao ⋅ Anshul Thakur ⋅ David A Clifton
Abstract
Large language models (LLMs) often produce confident but incorrect answers in settings where abstention would be safer. Standard evaluation protocols, however, typically require a response and do not account for how confidence should guide decisions under different risk preferences, providing limited insight into whether a model's confidence estimates support reliable decision making. To address this gap, we introduce the Behavioral Alignment Score (BAS), a decision-theoretic, black-box metric for evaluating how well LLM confidence supports abstention-aware decision making. BAS assigns realized utility to model predictions under an explicit cost model and aggregates performance across a continuum of risk thresholds, yielding a unified measure of decision-level reliability that depends on both the magnitude and ordering of model confidence. We show theoretically that truthful confidence estimates uniquely maximize expected BAS utility, linking calibration to decision-optimal behavior. Using BAS alongside standard metrics such as ECE and AURC, we then construct a comprehensive benchmark of self-reported confidence across multiple LLMs, tasks, and elicitation methods. Our results reveal substantial variation in decision-useful confidence, and while larger and more accurate models tend to achieve higher BAS, even frontier models remain prone to severe overconfidence in challenging settings. Importantly, models with similar ECE or AURC can exhibit very different BAS due to rare but highly overconfident errors, highlighting limitations of standard metrics. We further show that simple interventions, such as top-$k$ elicitation and post-hoc calibration, can meaningfully improve confidence reliability. Overall, our work provides a principled metric and an empirical benchmark for assessing the reliability of confidence estimates from black-box LLMs.
Successful Page Load