A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline
Abstract
Agentic AI tools offer a promising path to automating software development bottlenecks in scientific research pipelines: stages that take domain experts days to months to build, but where scientists care about correctness and robustness, not implementation details. We present an empirical study of general-purpose coding agents on a fly optogenetics data-to-discovery pipeline, with tasks substantially larger than existing benchmarks, datasets orders of magnitude bigger, and evaluation criteria grounded in domain expert standards. Agents can solve several individual pipeline stages, suggesting stage-level automation is tractable, but struggle with end-to-end tasks in ways that reveal a qualitatively different failure mode. We identify challenges largely absent from existing benchmarks, including computational resource management and generalization to large held-out data collections, and find that agents rarely exhibit the exploratory, visually-grounded validation behavior scientists use to diagnose subtle failures, a key open challenge. Finally, we distill principles for constructing rigorous evaluation criteria for open-ended problems.