TLDRocket
Sign in

Predictive Human Preference: From Model Ranking to Model Routing

Chip Huyen

A researcher developed a preference predictor that forecasts which AI model users will prefer for specific prompts by training on 20,927 comparison matches from LMSYS's Chatbot Arena dataset. The predictor achieved 76.2% accuracy when incorporating prompts as input, compared to 74.1% accuracy from Chatbot Arena's Bradley-Terry ranking that only considers model pairs. This enables model routing to direct queries to cheaper or faster models when they perform comparably to stronger models, potentially reducing costs and latency while maintaining response quality.

Why it matters

A challenge of building AI applications is choosing which model to use. What if we don’t have to? What if we can predict the best model for any prompt? Predictive human preference aims to predict which model users might prefer for a specific query. Human preference has emerged to be both the Northstar and a powerful tool for AI model development. Human preference guides post-training techniques including RLHF and DPO. Human preference is also used to rank AI models, as used by LMSYS’s Chatbot Arena. Chatbot Arena aims to determine which model is generally preferred. I wanted to see if it’s possible to predict which model is preferred for each query. One use case of predictive human preference is model routing. For example, if we know in advance that for a prompt, users will prefer Claude Instant’s response over GPT-4, and Claude Instant is cheaper/faster than GPT-4, we can route this prompt to Claude Instant. Model routing has the potential to increase response quality while reducing c

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.