HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning
Abstract
Self-reflection is a powerful mechanism for improving LLM reasoning, yet its effectiveness hinges on accurately attributing failures to specific reasoning steps---a capability that current models notably lack. Existing failure attribution methods either require expensive step-by-step counterfactual testing that scales poorly with trajectory length, or treat reasoning traces as flat sequences that ignore non-linear logical dependencies. We propose a hypergraph-based paired failure attribution (HPFA) framework, where analysis reason about the failure root cause through comparing the hyperedges of the targeted failure reasoning path against a reference successful path. By reducing the search space, our method efficiently localizes root causes and enables scalable synthesis of attribution data for training a lightweight attributor model via supervised fine-tuning and reinforcement learning. Experiments on math and coding tasks demonstrate that HPFA can dramatically increase the attribution accuracy and efficiency, and the trained attributor consistently improves reasoning accuracy at test time, outperforming baselines that lack graph structure or paired analysis.