PAM: Training Moderation Filters from Policy-Derived Supervision
Abstract
Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet existing filters often focus narrowly on safety, falling short of the broader alignment needs seen in real-world deployments. We introduce Policy-Aligned Moderation (PAM), a framework that compiles natural language policies into structured supervision, in the form of labeled prompt–response pairs with compliance scores, enabling the training of custom moderation filters without manual annotation. By transforming policies into training signals rather than relying on inference-time reasoning, PAM supports scalable, application-specific alignment across diverse policy settings. PAM-trained filters match or exceed strong safety filters and policy reasoning models on public safety benchmarks, and outperform them on PAMBENCH, four newly introduced human-annotated policy enforcement benchmarks targeting age restrictions, dietary accommodations, cultural alignment, and limitations in medical guidance. These performance gains are achieved while the PAM filter runs 5–100× faster at inference than policy-conditioned reasoning models.