KTPO: K-Step Test-Time Policy Optimization for Long- Horizon Discovery
Abstract
Large Language Models (LLMs) are emerging as a promising approach for discovery problems in mathematics and systems optimization. Existing search algorithms keep models frozen, yet applying reinforcement learning (RL) over full search trajectories spanning hundreds of refinement steps remains difficult. We present KTPO (K-Step Test-Time Policy Optimization), which addresses this challenge by training on short trajectories of k refinement steps. To retain the benefits of longer search, each trajectory starts from a previously discovered program rather than from scratch. In particular, KTPO selects programs near the model’s current capability frontier, estimated by the model’s improvement rate over existing programs. KTPO with 4B-9B models achieves strong performance across four mathematical and systems optimization tasks. On mathematical problems, KTPO outperforms or approaches the human record on Circle Packing and the Erd˝os minimum overlap problem. On systems optimization tasks from ADRS, a benchmark suite for real-world systems optimization, KTPO outperforms human record on maximizing SQL query prefix hit rate, and surpasses the prior AI record on minimizing multi-cloud data transfer cost.