9 Judges, Really 2: Correlated Errors Undermine LLM Evaluation Panels
Key point
A study found that because models in LLM evaluation panels repeat similar errors, adding multiple models to the vote provides very little actual information value.
Details
The LLM-as-a-judge approach aims to increase evaluation reliability by aggregating votes from multiple models. However, research found that a panel using 9 frontier LLMs drawn from 7 model families actually provides only about 2 models' worth of independent information value (Kish effective sample size).
The main cause is Correlated Errors, where models repeat the same mistakes on the same items. This causes about three-quarters of the panel's nominal independence to be lost, and the panel's actual accuracy is measured 8-22 percentage points lower than would be expected if the votes were independent.
The study's key findings are as follows:
- Single-model superiority: Under all conditions, the single best-performing model performed better than or on par with the entire panel.
- Scaling limits: Increasing the number of judge models or using more sophisticated aggregation algorithms had limited effect in closing this gap.
- Robustness confirmed: This phenomenon was consistently observed across prompt variations, temperature settings, Chain-of-Thought (CoT) reasoning, and pairwise preference tasks such as RewardBench.
In conclusion, the bottleneck in LLM evaluation lies not in the aggregation algorithm but in the correlation among judge models, and simply increasing the number of models cannot substitute for genuinely independent evaluation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.