Recovering Wasted Compute in Autoresearch Agents
Abstract
A slew of recent works have developed agents for solving data science tasks end-to-end. Such agents have inspired large industry investment, motivated by their potential to automate time-consuming human labor and customize machine learning solutions for specialized applications. In this paper, we identify common failure modes of these agents when applied to tabular datasets: buggy code generation, absent hyperparameter tuning, inefficient usage of compute budget, and insufficent exploration. We further explore patches for these failure modes to enable efficient and effective solution generation. In particular, we find that prompt and control level enhancements, a debug consultant that shares discovered runtime constraints across all branches of the search tree, and refined tree search algorithms successfully mitigate failure modes. Our results indicate that significant gains in data science agent performance are achievable through better agentic design alone, independent of the underlying language model's capabilities.