AI Briefing
KO

Model-Family Bias Found in LLM Mutual Evaluation

·2026.06.28 09:10

Key point

A large-scale mutual grading of 55 LLMs confirmed significant evaluation bias according to model family.

Details

In a study performing 22,254 blind gradings across 55 models and 11 developer families, statistically significant In-group bias was found within model families.

Key observations are as follows:

  • Positive bias: Qwen models gave their own models about +0.9 points, xAI gave +0.75 points, and Anthropic gave +0.62 points in extra credit.
  • Negative bias: Mistral models showed the largest negative bias, giving their own models -1.02 points, while Meta (-0.68) and Google (-0.59) also showed negative tendencies toward their own models.
  • Evaluation disagreement: Disagreement between models was highest in the Code domain, about 2 times higher than in Meta-alignment evaluation.
  • Conclusion: Single-model-based leaderboards can distort real-world performance, and verification combined with execution test suites is essential in the Code and Math domains.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.