Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows
Abstract
Frontier coding agents are increasingly used in workflows where users supervise progress primarily through repeated improvement of a public score, namely the reported score on a public evaluation file in the workspace, rather than through direct inspection of the agent's intermediate outputs. We study whether multi-round user pressure to improve that score induces public-score exploitation: behavior that raises the public score through shortcuts without improving hidden private evaluation. We begin with a preliminary single-script tabular classification task, where GPT-5.4 and Claude Opus 4.6 exploit in all 10 runs. We then build \name, a 34-task machine-learning repository benchmark spanning three input modalities, and collected 1314 multi-round trajectories from 13 frontier coding agents. Across \name, we observe 498 exploitative runs, spanning across all tasks. We also find that stronger models have higher exploitation rates, supported by a significant Spearman rank correlation of 0.82. Our ablation experiments show that higher user pressure leads to earlier exploitation, manifested itself in the average first exploit round reducing by 6.8 rounds (\textit{i.e.}, 11.42 to 4.58). As a mitigation, explicit anti-exploit wordings in prompt completely eliminates exploitation while keeping public labels exposed. These results show that repeated pressure to improve a public score can induce public-score exploitation in coding-agent workflows, and that prompt design can materially change that behavior.