Overton Pluralistic Reinforcement Learning for Large Language Models
Yu Fu ⋅ Seongho Son ⋅ Ilija Bogunovic
Abstract
Existing alignment paradigms remain limited in capturing the pluralistic nature of human values. Overton Pluralism (OP) addresses this gap by generating responses with diverse perspectives from a single query. This paper introduces OP-GRPO (Overton Pluralistic Group Relative Policy Optimization), a reinforcement learning framework for implicit Overton Pluralism that enables a single LLM to produce pluralistic responses without explicit prompting or modular orchestration. Our workflow consists of two main steps: (1) similarity estimator training, which fine-tunes a Sentence Transformer for OP tasks to provide more accurate coverage evaluation of generated responses; and (2) OP-GRPO training, which incorporates this similarity estimator into a carefully designed dual-reward system to ensure both broad coverage of genuine human perspectives and the uniqueness of each perspective, thereby promoting diversity. Empirical results demonstrate a “Small models, Big perspective coverage” effect: the trained $\texttt{Qwen2.5-3B-Instruct}$ surpasses the $\texttt{GPT-OSS}$ (20B) baseline with a $37.4\%$ relative accuracy gain on the Natural Language Inference (NLI) benchmark. It also outperforms a modular-architecture baseline by $19.1\%$. Evaluations with GPT-4.1 as LLM judge and OP-Benchmark further confirm the robustness of our approach.
Successful Page Load