UNMASK Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
Abstract
Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on adversarial or out-of-distribution inputs. Existing approaches either require manual specification of the feature vocabulary or automate discovery only partially, leaving the gap between dataset-level correlation and model-level exploitation unaddressed. We present UNMASK, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without human annotation of spurious features. Given unlabeled training examples, UNMASK generates candidate surface patterns as executable boolean expressions, filters them through a multi-stage statistical validation protocol with independent replication, and establishes causal model dependence via programmatically verified counterfactual interventions. Causally confirmed features then serve as annotation-free group definitions for Deep Feature Reweighting, eliminating the group labels that standard DFR requires. Applied to BERT-base-uncased and RoBERTa-base trained on MNLI, our pipeline independently rediscovers established lexical-overlap and negation biases, reducing causal reliance by 60-62% on verified features and improving adversarial robustness by up to 4.6 pp on ANLI and 9.2 pp on HANS, with DFR-IID yielding the strongest20 out-of-distribution gains for RoBERTa-base. We further demonstrate that the discovery and validation stages generalize to reward model preference data, surfacing interpretable spurious correlations in RewardBench2, indicating broader applicability across natural language datasets.