AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Predictive Human Preference: From Model Ranking to Model Routing

Collected Oct 1, 2026

Chip Huyen published an experiment on predictive human preference: predicting which model users would prefer for a specific query, rather than only ranking models overall. The stated use cases are model routing and interpretability.

She evaluated the correctness of Chatbot Arena's ranking using LMSYS's July 2023 dataset of 33,000 crowd-sourced comparisons among 20 models, which includes the prompt for each match. Chatbot Arena switched in December 2023 from Elo to Bradley-Terry, scaling the scores to look Elo-like. In a 10% test set of 3,300 examples, the model with the higher Bradley-Terry score was preferred 74.1% of the time on the 2,367 non-tie matches, and 85.1% on the 355 non-tie matches involving GPT-4. GPT-4's Bradley-Terry score in this dataset was 1189, ahead of Claude-v1's 1150 and Claude-instant-v1's 1110.

Huyen treated preference prediction as binary classification, taking (prompt, model_a, model_b) as input, using DistilBERT as a prompt encoder and training on 90% of the dataset. Non-tie matches only performed better than including ties. On the same test set, the predictor reached 75% accuracy without prompts and 76.2% with prompts, versus Chatbot Arena's 74.1%; for GPT-4 matches it reached 86.2% and 87% versus 85.1%. She noted the predictor was trained on a small amount of noisy crowd-sourced data, and that among the 33,000 prompts, 180 (0.55%) were "hello" or "hi".

Using all 190 model pairs per prompt and a Bradley-Terry model, she generated per-prompt leaderboards, which showed that GPT-4 led for the prompt "Derive the elastic wave equation" with a score of 1214, and that score spread was smaller for a simple prompt than a challenging one. Huyen named Martian, which announced a $9M seed round, as working on model routing, said LMSYS is also working on it, and noted four groups total including two in stealth. She said LMSYS told her using GPT-4 to compare responses works better than noisy crowd-sourced annotations, estimating 10,000 comparisons would cost $200-500. She credited Luke Metz for experiment help and Han-chung Lee for plot feedback.

Read at Chip Huyen

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

A challenge of building AI applications is choosing which model to use. What if we don’t have to? What if we can predict the best model for any prompt? Predictive human preference aims to predict which model users might prefer for a specific query. Human preference has emerged to be both the Northstar and a powerful tool for AI model development. Human preference guides post-training techniques including RLHF and DPO . Human preference is also used to rank AI models, as used by LMSYS’s Chatbot Arena . Chatbot Arena