No Single Best Model for Diversity: Learning a Router for Sample Diversity
Yuhan Liu ⋅ Fangyuan Xu ⋅ Vishakh Padmakumar ⋅ Daphne Ippolito ⋅ Eunsol Choi
Abstract
When posed with prompts that permit a large number of valid answers, comprehensively generating them is the first step towards satisfying a wide range of users. In this paper, we study methods to elicit a comprehensive set of valid responses. To evaluate this, we introduce **diversity coverage**, a metric that measures the sum of quality scores assigned to each **unique** answer in the predicted answer set relative to the best possible answer sets with the same number of answers. Using this metric, we evaluate $18$ LLMs, finding no single model dominates at generating diverse responses to a wide range of open-ended prompts. Yet, per each prompt, there exists a model that outperforms all other models significantly $70$% of the times. Motivated by this, we introduce a router that predicts the best model for each query. On *WildChat* our trained router outperforms the single best model baseline ($26.3$% vs $23.8$%). We further show generalization to an out-of-domain dataset (*NoveltyBench*) as well as different answer-generation prompting strategies. Our work lays foundation for studying generating comprehensive answers when we have access to a suite of models.
Successful Page Load