AI Briefing
KO

Disagreement Among Frontier LLMs in Real-World Fact-Checking

·2026.05.29 09:33

Key point

Five frontier LLMs were found to reach different verdicts in 67% of cases when fact-checking real-world claims.

Details

Analyzing 5 frontier LLMs on 1,000 real user claims, models disagreed on the verdict in 67% of cases. All models agreed on the same verdict in only 33% of cases.

Key findings:

  • Substantial disagreement: Using a 4-level rubric (True, Mostly True, Misleading, False), 'substantial disagreement'—defined as a gap of 2 or more levels—occurred in 34% of cases, and extreme splits between True and False reached 21%.
  • Pairwise model agreement: Label agreement rates between model pairs ranged from 53% to 75%. The agreement rate was highest, at 75%, between Gemini 3 Pro and its Search version, which share the same base model.
  • Model-specific verdict tendencies: Clear differences in tendencies emerged across models. Gemini 3 Pro had a very high share of True verdicts (54%), while Claude Opus 4.7 tended to use the intermediate levels (Mostly True, Misleading) more frequently.

By analyzing the structure of disagreement between models without ground-truth labels, this study focused on revealing the structural instability of verdicts rather than the accuracy of the LLMs themselves.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.