AMAT: Automated Multi-Agent Topology Design via Reinforcement Learning
Abstract
Large Language Model-based multi-agent systems have demonstrated remarkable capabilities in handling complex tasks through collective collaboration. However, current design approaches rely on either manual crafting or automated heuristic search, yet suffer from two critical limitations: ephemeral experience and value-agnostic optimization. To address these fundamental challenges, we propose AMAT, a self-bootstrapping reinforcement learning framework that shifts the paradigm from case-by-case online search to persistent offline policy learning. AMAT formulates the topology design of multi-agent systems as a constrained Markov decision process on graphs, and integrates Monte Carlo Tree Search with a policy-value network to internalize structural patterns and value judgments. Comprehensive evaluations across five benchmarks demonstrate that AMAT (1) consistently outperforms both handcrafted and automated multi-agent baselines across all evaluated benchmarks, (2) substantially reduces both training expenditure and inference latency compared to existing search-based approaches, and (3) exhibits strong cross-task generalization within task domains and cross-LLM-backbone transferability.