HAI-Agent: Improving Long Horizon Software Engineering with Handoff Interventions
Abstract
Multi-agent systems for software engineering often suffer from a \textit{delegation tax}, where the coordination overhead of sub-agents degrades performance compared to single-agent baselines. We trace this penalty to failures at the inter-agent \textit{handoff boundary}, defined by \textit{context loss} during task delegation and \textit{weak verification} upon return. We propose HAI-Agent, a framework that transforms this boundary into a governed interface through three \textit{handoff interventions}: (1) \textit{Delegation Review}, a pre-dispatch gate ensuring task specifications are self-contained; (2) \textit{Retrospective Reflection}, a mandatory self-certification by the agent; and (3) \textit{Verification Fork}, an independent validation of the return by the orchestrator. Our experiments demonstrate that \ours{} substantially outperforms existing specialized scaffolds across multiple model scales. On SWE-bench-Pro, HAI-Agent achieves a \textit{46.8\%} resolve rate with Qwen3.5-397B-A17B. On Commit0-lite, it reaches a \textit{64.0\%} pass rate with Qwen3.5-397B-A17B and \textit{61.3\%} with Claude-4.5-Sonnet, representing a \textit{+10.8\%} absolute gain over the SWE-Agent. Our results prove that governing the handoff transforms delegation from a performance liability into a reliable multiplier for software engineering.