AI Briefing
KO

LLM Performance and Agent Effectiveness Examined Through IMO Math Problems

·2026.07.26 16:21

Key point

Using IMO math problems, the reasoning performance of frontier models and open-weight models and the effectiveness of multi-agent usage were compared.

Details

This is a result of testing the reasoning abilities of various LLMs using International Mathematical Olympiad (IMO) problems as a benchmark. IMO problems are new problems not included in training data, and since they require complex multi-step reasoning, they are suitable for measuring a model's general intelligence.

Key Results:

  • Frontier Models: The Sol and Fable models recorded perfect or near-perfect scores even without a separate harness.
  • Claude Models: Sonnet and Opus showed low performance in the basic web app, but their performance improved significantly through provider harnesses such as Claude Code and AutoFyn, a self-developed multi-agent harness.
  • Open-weight Models: GLM showed a level similar to Sonnet without a harness, and its performance improved when AutoFyn was applied.

Limitations and Observations:

  • Hallucination Problem: Even in mathematics, a verifiable domain, cases were found where models claimed incorrect solutions.
  • Limits of Reasoning: For the most difficult problem (P3), even frontier models with all harnesses applied missed the key idea (Key reduction) and got stuck at the same step. This suggests the need for core intuition that goes beyond simple search or verification.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.