Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model
Yongmin Kim ⋅ Shota Takashiro ⋅ Yusuke Iwasawa ⋅ Takeshi Kojima ⋅ Yutaka Matsuo
Abstract
Large Reasoning Models (LRMs) achieve strong performance on complex tasks through extended chain-of-thought generation, but incur substantial computational costs during inference. In production settings, batched inference is essential for high throughput, yet existing adaptive pruning methods face performance limitations: First, they rely on averaging activations across samples to determine shared pruning masks, which may miss critically activated neurons for some individual samples. Second, after the averaging, they adopt threshold-based selection for pruning neurons, causing sparsity ratio instability. In this work, we propose a training-free adaptive pruning method designed specifically for batched inference in LRMs. Since averaging activations across samples can miss neurons that are critical for individual samples, our method adopts max-pooling for cross-sample aggregation to preserve such sample-specific important neurons. To stabilize the activation sparsity ratio, our method adopts periodic top-k selection over the aggregated neurons instead of threshold-based selection. Furthermore, based on the observation that important neurons tend to be repeatedly activated, we incorporate an activation memory mechanism to capture periodically important neurons. Experiments on diverse reasoning benchmarks demonstrate that our method outperforms the previous state-of-the-art adaptive pruning method by 39.7 percentage points in average accuracy at batch size 4 with 50\% target sparsity on DeepSeek-R1-Distill-Qwen-7B, along with $1.40\times$ speedup over dense inference at 50\% actual sparsity, demonstrating practical efficiency gains for deployment.
Successful Page Load