RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback
Abstract
With the rapid advancement of Large Language Models (LLMs), developing effective critic modules for precise guidance has become crucial yet challenging. An ideal critic should both judge accurately and drive refinement, yet existing approaches focus primarily on judgment, leaving the capacity to guide refinement underexplored. To address this gap, we propose RefCritic, a long-chain-of-thought critic trained via reinforcement learning with dual rule-based rewards: (1) instance-level correctness of solution judgments and (2) gains of policy-model refinements driven by the critique, encouraging evaluations that are both reliable and actionable. Extensive experiments on Qwen2.5-14B-Instruct and DeepSeek-R1-Distill-Qwen-14B demonstrate that RefCritic achieves high judgment accuracy and surpasses step-level supervised baselines on fine-grained error localization with only solution-level supervision. Moreover, RefCritic produces meaningful evaluations that substantially benefit refinement, achieving +6.8\% and +9.3\% gains on AIME24 and AIME25 after one round of critique-refinement, and exhibits even stronger gains +13\% and +10\% as the computational cost scales up, indicating seamless adaptability to test-time strategies. RefCritic establishes that optimizing critics for refinement effectiveness, rather than solely for judgment, can bridge the gap between evaluation quality and practical utility.