RankBALD: Ranking-Aligned Active Evaluation for Language Models
Abstract
Language model evaluation faces a structural challenge: as performance gaps between state-of-the-art models narrow, static evaluation sets require increasingly large budgets to yield stable, reliable model comparisons. Adaptive evaluation can improve measurement efficiency by sequentially selecting informative test items, but existing approaches primarily focus on estimating latent parameters that describe model capabilities. This focus overlooks the fact that experimenters may have diverse downstream evaluation goals, such as to produce a relative ordering of models. We introduce a framework for adaptive evaluation in the Item Response Theory (IRT) setting that is compatible with both capability assessment and model-ranking objectives. Within this framework, ranking reliability depends on resolving uncertainty near the boundaries of pairwise orderings rather than minimizing overall posterior variance. To operationalize this insight, we draw on Bayesian active learning (BAL), a family of methods that select examples by maximizing expected information gain under a posterior belief. BAL was originally developed for information-efficient parameter identification. We adapt BAL to the IRT evaluation setting, giving rise to what we call \textit{active evaluation}: sequential item selection driven by principled uncertainty reduction over model abilities and their relative ordering. We develop a family of acquisition criteria that increasingly align with the ranking objective, culminating in \textbf{RankBALD}, which directly maximizes the mutual information between candidate test items and the pairwise model-ordering variables. Experiments across six English benchmarks from the Open LLM Leaderboard and five African-language benchmarks show that, compared with ability-estimation baselines, ranking-aligned acquisition consistently improves ranking validity and reduces evaluation variance. These results demonstrate that aligning item selection with the downstream evaluation objective substantially improves the efficiency and reliability of active evaluation.