Fork-think with Confidence
Abstract
Parallel thinking has enjoyed large success for boosting LLM performance on reasoning tasks without the need for any re-training. However, existing methods first sample multiple reasoning paths--which inevitably leads to overgeneration-- and then apply pruning or early stopping to compensate. Instead, first identifying points to sample from has the potential to avoid such unnecessary generations, but has been underexplored so far. Based on this observation, we propose Fork-think with Confidence, where we first identify forking points using a single reasoning path and then trigger thinking, i.e., sample multiple continuations and aggregate them for the final response. Our experiments across three challenging reasoning benchmarks and three model families show that Fork-think reduces the token consumption by up to 30% and run-time by up to 57\%, while matching (or surpassing) the performance of parallel thinking. Moreover, Fork-think performs competitively against state-of-the-art pruning and early stopping techniques and does not need any warm-up or offline training, establishing pre-determined forking as a promising paradigm for efficient LLM reasoning.