PrefPO: Pairwise Preference Prompt Optimization
Rahul Singhal ⋅ Pradyumna Tambwekar ⋅ Karime Maamari
Abstract
Prompt engineering is effective but labor-intensive, motivating automated optimization methods. Existing methods produce verbose, repetitive prompts and typically require labeled datasets that are often unavailable. We introduce PrefPO, a lightweight prompt optimization approach inspired by reinforcement learning from human feedback (RLHF). Its preference-based approach reduces the need for labeled data and hyperparameter tuning—only a starting prompt and natural language criteria are needed. PrefPO uses an LLM discriminator to express pairwise preferences over model outputs and provide feedback to an LLM optimizer, iteratively improving performance. We evaluate PrefPO on 9 BIG-Bench Hard tasks and IFEval-Hard, a newly-curated, challenging subset of IFEval. PrefPO matches or exceeds SOTA methods, including GEPA, MIPRO, and TextGrad, on $6/9$ tasks and performs comparably to TextGrad on IFEval-Hard ($82.4\\%$ vs $84.5\\%$). Unlike other methods, PrefPO can optimize in both labeled and unlabeled settings. Without labels, PrefPO matches its labeled performance on $6/9$ tasks, proving effective without ground truth. PrefPO also improves *prompt hygiene*: we find existing methods produce prompts $14.7$x their original length or with $34\\%$ repetitive content; PrefPO reduces these issues by $3$–$5$x. Furthermore, both LLM ratings and human preferences favor PrefPO's prompts over TextGrad's. Finally, we identify *prompt hacking* in prompt optimizers, where methods game evaluation criteria, and find PrefPO is susceptible at nearly half the rate of TextGrad ($38\\%$ vs $66\\%$), generating fewer brittle, misaligned prompts.
Successful Page Load