Data Canvas: A Provenance-Guided Harness for Agentic Data Engineering
Abstract
As agentic LLM-based technologies mature, there is tremendous promise in using them to design and execute complex data engineering tasks, compiling and integrating specialized datasets on behalf of users. These tasks span from high-level planning with large reasoning models and query optimization, to lower-level runtime components for information extraction and semantic reasoning. Recent frontier systems show strong potential on such tasksbut they remain prone to hallucinations, omissions, and underspecified decisions that are difficult to inspect or correct. We argue that agentic data engineering needs a provenance-guided harness: a control layer around planning and execution that makes outputs attributable, inspectable, and steerable through feedback. We present Data Canvas, a provenance-guided harness for agentic data engineering that wraps LLM-driven execution with structured semantic operators, explicit intermediate tables, and a provenance graph over operator events and results. This harness supports sparse human or automated feedback by tracing errors to their responsible reasoning steps, propagating corrections to related outputs, and replaying only the affected portions of the workflow. Using standard deep research benchmarks, we show that Data Canvas improves answer quality over existing systems, and that provenance-guided feedback yields substantial gains at a small fraction of the cost of prompt re-engineering and re-planning.