Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models
Abstract
In high-stakes decision-making applications of large language models (LLMs), practitioners require not only accurate LLMs but also uncertainty estimates for their predictions. Existing approaches to uncertainty estimation for LLMs require access to log-probabilities output by the model or require fine-tuning access. However, the vast majority of industrial LLM products use closed-source API models, and many such API models like GPT-5 do not return log-probabilities and may not allow fine-tuning. We introduce PINOCCHIO, a suite of 8B checkpoints for assigning uncertainty estimates to predictions made by popular API models. To this end, we curate a large collection of hard datasets, ranging from long-context to vision-language, that expose conditions under which frontier models output incorrect responses. We then train PINOCCHIO to recognize the patterns in prompts and model responses that indicate confidence or uncertainty. PINOCCHIO checkpoints significantly outperform existing baselines for black-box uncertainty estimation in calibration metrics and transfer to unseen model families, including OpenAI, Qwen, Meta, Anthropic, and Google, whose responses they were not trained on. We release code for adding uncertainty estimation to existing repos in only two additional lines of code