EvoSkill: Automated Skill Discovery for Multi-Agent Systems
Salaheddin Alzubi ⋅ Noah Provenzano ⋅ Jaydon Bingham ⋅ Christian A Calvo ⋅ Weiyuan Chen ⋅ Tu Vu
Abstract
Coding agents are increasingly used as general-purpose problem solvers, but their flexibility does not by itself confer the domain expertise needed for specialized tasks. Recent work addresses this through $agent$ $skills$: reusable workflows and code that augment agents with domain-specific capabilities. Most skills today are hand-crafted, and existing evolutionary approaches optimize low-level artifacts (e.g. prompts and code) that are tightly coupled to specific models and tasks. We introduce **EvoSkill**, a self-evolving framework that automatically discovers and refines agent skills through iterative failure analysis. EvoSkill analyzes execution failures, proposes new skills or edits to existing ones, and materializes them into structured, reusable skill folders. A frontier of agent programs governs selection, retaining only skills that improve held-out validation performance while the underlying model remains frozen. We evaluate EvoSkill across four benchmarks (OfficeQA, SealQA, LiveCodeBench, and FRAMES) and multiple model families including Claude, GPT-OSS, Kimi, and Gemini, achieving gains of up to $+15.7\%$ on SealQA and $+11.4\%$ on OfficeQA. EvoSkill outperforms GEPA, a prompt-level evolutionary baseline. We further show that evolved skills transfer zero-shot across both models and datasets: skills evolved on SealQA improve BrowseComp accuracy by $+5.3\%$ without modification, and skills evolved with one model can improve performance when applied to a different model family.
Successful Page Load