AI Briefing
KO

LMArena Introduces Factuality Metric and Announces Rankings

·2026.07.16 01:49

Key point

LMArena has introduced a new ranking system that combines human preference and factuality.

Details

LMArena has introduced a new model ranking system that weights and combines human preference with Factuality scores. The Factuality metric is available as a selectable option in both Text Arena and Search Arena.

The measurement method works by extracting web-verifiable claims from randomly sampled responses in Arena battles and comparing the average accuracy of those claims. To date, more than 2 million claims have been labeled for this purpose.

Key ranking changes:

  • Text Arena (with Factuality applied): Claude Fable 5 fell slightly to 2nd place, while GPT-5.5 rose 13 spots to 7th place. Meanwhile, Muse Spark plunged from 7th to 20th place.
  • Search Arena: GPT-5.5-search took 1st place on the factuality leaderboard. GPT-5.2-search jumped significantly from 11th to 3rd place.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.