Fragility Under Pressure: Evaluating the Iterative Stability of LLMs in Constrained Interactive Coding
Abstract
Interactive coding has become an important mode of software development, where humans and large language models (LLMs) iteratively refine code under evolving requirements. Unlike static code generation, this setting requires models to preserve functional correctness while continuously incorporating new constraints across multiple turns. However, existing evaluations are largely single-turn and therefore provide limited insight into how models behave under iterative constraint pressure. To study this gap, we introduce STRIDE, a dynamic evaluation framework that simulates iterative code refinement with accumulating, verifiable constraints. Using STRIDE, we evaluate a broad set of state-of-the-art models and find that strong single-turn coding performance does not reliably translate into stable multi-turn behavior. In particular, models show a consistent asymmetry between different types of constraints: they handle avoidance-style requirements more robustly than prescriptive structural modifications. We also observe a recurring trade-off between surface-level compliance and functional correctness, where satisfying newly introduced constraints can degrade the underlying program logic. These results suggest that current LLMs remain brittle in interactive coding settings and that single-turn benchmarks may substantially overestimate their reliability for real-world collaborative software engineering.