Compared to What? Baselines and Metrics for Counterfactual Prompting
Abstract
Counterfactual prompting (i.e., perturbing a single factor and measuring output change) is widely used to evaluate things like LLM bias and CoT faithfulness. But in this work we argue that observed effects cannot be reliably attributed to the targeted factor without accounting for baseline “meaning-preserving” modifications to text that establish general model sensitivity. In the parlance of causal inference, this is because every counterfactual edit is a compound treatment that bundles the variable of interest with incidental surface-form variation; this violates treatment variation irrelevance. We observe prediction flip rates on MedQA of 14.9% when we surgically change patient gender. However, this is statistically indistinguishable from the flip rates induced by simply paraphrasing inputs (14.1%). In this case, it would therefore be unwarranted to conclude that the LLM is especially sensitive to patient gender. To account for this and robustly measure the effects of targeted interventions, we propose a framework in which we compare (via statistical testing) differences observed under target interventions to those induced by paraphrasing inputs (adjusted to match perturbation token edit lengths). We then use this framework to revisit a prior analysis done on the MedPerturb dataset, which reported evidence of model sensitivity to patient demographics and stylistic cues. We find that these effects largely dissipate when we account for general model sensitivity, with only 5 of 120 tests reaching statistical significance; this calls into questions prior claims about models being sensitive to these factors. Applying the same framework to occupational biography classification (Bias-in-Bios), we detect highly significant directional gender bias, demonstrating that the framework identifies real directional effects even when they are small. We evaluate a range of metrics---aggregate, per-sample distributional, and regression---and find that per-sample metrics (JSD, KL) are dramatically more powerful than aggregate metrics (MI, ɸ, flip rate) and regression powerfully and uniquely characterizes effect direction and magnitude. We provide empirical guidance on running counterfactual prompting experiments for LLMs.