What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
Abstract
In domains like medicine and finance, large-scale labeled data is costly and often unavailable, leading to models trained on small datasets that struggle to generalize to real-world populations. Large language models contain extensive knowledge from years of research across these domains. % Existing methods either prompt LLMs to generate stochastic priors or directly predict labels, both introducing variability. Moreover, the performance of these methods on out-of-distribution datasets remain unexplored. We propose LoID (Logit-Informed Distributions), a deterministic method for extracting informative prior distributions for Bayesian logistic regression by directly accessing their token-level predictions. Rather than relying on generated text, we probe the model's confidence in opposing semantic directions (positive vs. negative impact) through carefully constructed sentences. By measuring how consistently the LLM favors one direction across diverse phrasings, we extract the strength and reliability of the model's belief about each feature's influence. We evaluate LoID on ten real-world tabular datasets under synthetic out-of-distribution (OOD) settings characterized by covariate shift, where the training data represents only a subset of the population. We compare our approach against (1) standard uninformative priors, (2) AutoElicit, a method that prompts LLMs to generate expert-style priors via text completions, (3) LLMProcesses, a method that prompts LLMs to generate predictions and (4) an oracle-style upper bound derived from fitting logistic regression on an in-distribution training set. We assess performance using Area Under the Curve (AUC). Across datasets, LoID significantly improves performance over logistic regression trained on OOD data, recovering up to \textbf{59\%} of the performance gap relative to the oracle model. LoID outperforms AutoElicit and LLMProcessesc on 17 out of 20 datasets, while providing a reproducible and computationally efficient mechanism for integrating LLM knowledge into Bayesian inference.