A Diagnostic Failure Taxonomy and Trajectory Critic for Data Agent Improvement
Suchen Liu ⋅ Yuanfeng SONG ⋅ Jun Gao ⋅ Xing Chen
Abstract
Large Language Models (LLMs) have shown strong potential as autonomous agents for complex data analysis. However, extending their capabilities to open-ended insight discovery remains challenging, as failures inevitably arise during multi-step reasoning and often cascade into severe downstream failures. Existing generic self-correction methods struggle to address these issues due to a lack of structural understanding of analytical failures. To tackle this, we propose the \textbf{Taxonomy-Guided Trajectory Critic (TGT-Critic)}, a novel runtime correction framework that detects and mitigates failures within the agent's execution trajectory. To ground our critic with domain-specific diagnostic capabilities, we first construct a fine-grained Agent Failure Taxonomy and an annotated dataset (InsightBench) to systematically map the root causes of analytical failures across six dimensions. By formalizing failure diagnosis as a taxonomy-grounded inference problem, TGT-Critic dynamically provides targeted revisions for both high-level planning and low-level action execution. Experimental results demonstrate that our framework effectively arrests cascading failures, significantly outperforming generic baselines and yielding a 10.52$\%$ relative gain in the quality of generated insights.
Successful Page Load