AivexaNewsSearch
AI news for builders and product teamsChecked every hour

When LLM judges agree, should we believe them?

Collected Oct 1, 2026

Amazon Science researchers have introduced a method for aggregating judgments from multiple large language models that accounts for correlations among the judges, according to a paper presented at this year's International Conference on Machine Learning (ICML). The paper, "Dependence-aware label aggregation for LLM-as-a-judge via Ising models," is coauthored with Shiva Kasiviswanathan.

The work addresses a limitation in common aggregation approaches such as uniform and weighted majority vote, which assume judges make errors independently. The authors write that judges may share a prompt template, a training lineage, a model family, or a common blind spot, in which case agreement may reflect repeated mistakes rather than independent evidence.

The proposed method treats a judge panel as a network and models pairwise dependence between binary judge outputs with an Ising model. It is designed for the unsupervised setting, learning from judge outputs without human reference labels, and treats each item's true label as a latent variable inferred jointly with parameters for judge reliability and dependence. Two variants are described: one in which the relationship pattern is treated as roughly the same for positive and negative labels, and a class-dependent model that lets the pattern change with the label, which requires more data.

Evaluation covered three binary tasks: relevance classification for retrieved information, toxicity classification, and summarization assessment. The judge panel contained 10 judge models run at temperature zero. Using all 10 judges and maximum available training data per task, the strongest dependence-aware results were 0.912 accuracy on relevance versus 0.820 for weighted majority vote and 0.804 for uniform majority vote; 0.792 on toxicity versus 0.694 and 0.695; and 0.806 on summarization versus 0.737 and 0.561.

The authors write that modeling dependence improved accuracy once the system had enough evaluation items and judges to estimate meaningful relationships, and state that the method outperformed the best-performing baseline — a panel weighted according to historical accuracy — by 9% to 14% on standard metrics. They suggest practices including evaluating the panel rather than only individual judges, treating model diversity as statistical diversity, inspecting agreement structure, and reporting uncertainty with dependence in mind.

Read at Amazon Science

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Discounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.