FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
Abstract
Large language model agents increasingly solve complex tasks by constructing inference-time workflows that combine reasoning, tool use, and code execution. While such workflows enable flexible problem solving, the useful procedures discovered during execution are often transient: they help solve the current task but are not retained in a form that can systematically benefit future tasks. This paper presents FlowEvo, a training-free framework that turns successful workflows into validated executable skills and maintains them in a persistent skill bank at inference time. These accumulated skills influence future behavior in two complementary ways: they can be invoked directly when a new task closely matches prior experience, and they can also serve as structured references for constructing new workflows when direct reuse is insufficient. Through this workflow--skill--workflow feedback loop, FlowEvo enables agents to accumulate and refine task-solving capability over time without updating model parameters. Experiments on aligned subsets of HumanEval, MBPP, GSM8K, and MATH show that FlowEvo consistently outperforms strong workflow-optimization baselines. Further analyses on MBPP show gains both in exact skill-reuse settings and in more challenging settings where prior skills improve workflow construction without being replayed verbatim. The code is public at \url{https://anonymous.4open.science/r/FlowEvo-7043/}.