AI Briefing
KO

5 Frontier LLMs Give Different Answers on 67% of 1,000 Real-World Fact-Check Items

·2026.05.28 21:20

Key point

When 5 top-tier LLMs analyzed 1,000 real-world fact-check items, the models disagreed with each other on 67% of the items.

Details

According to research by Lenz Research, when 5 frontier LLMs reviewed 1,000 real user fact-check requests, the models failed to agree in 67% of cases.

The key findings are as follows:

  • 67% disagreement: In 672 out of 1,000 cases, there was either a model that diverged from the majority opinion, or no clear majority was formed at all.
  • 34% substantive gaps: In 343 cases, the models showed serious disagreement, with answer categories (True, Mostly True, Misleading, False) differing by 2 levels or more.
  • Limited level of agreement: Krippendorff's $\alpha$ (ordinal) came in at 0.639, indicating that while the models are not answering randomly, the agreement is not strong enough to serve as a single unified standard of judgment.

Looking at the models' judgment patterns, they tended to converge on clear-cut verdicts at both extremes, such as 'True' and 'False', but diverged significantly on intermediate categories such as 'Mostly True' and 'Misleading'. Additionally, some models concentrated their answers at the two extremes, while others showed a tendency to spread their answers broadly across the intermediate categories.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.