WorldPM: Scaling Human Preference Modeling via Real-World Feedback
Abstract
Inspired by the success of scaling language modeling through massive real-world corpora, we demonstrate that similar power-law patterns exist in preference modeling. We propose \textbf{World Preference Modeling (WorldPM)}, which model a robust consensus on intrinsic utility from real-world user interactions. We curate a 15M-scale preference dataset from online forum data and conduct large-scale experiments on models from 1.5B to 72B parameters. Across 10 objective preference test sets with consensus-based labels, we observe two distinct scaling patterns: (1) common-sense preference (e.g., identifying factual errors) exhibits stable log-linear improvement across all model sizes, while (2) domain preference (e.g., coding, math, knowledge QA) reveals an emergent phenomenon where only sufficiently large models achieve meaningful scaling. Through training dynamics analysis, we reveal the learning mechanism behind preference scaling: the model first learns to discriminate using surface preferences, but counter-surface examples progressively compel it to discover intrinsic utility beyond surface features as training scales. Further experiments validate WorldPM as an effective foundation for preference fine-tuning, broadly improving generalization across human preference datasets of varying sizes (7K, 100K, and 800K), with gains exceeding 5\% on many key subtasks.