AI Briefing
KO

Opus 4.8, ARC-AGI-3 Breakthrough (1-min read)

·2026.06.02 09:00

Key point

The latest LisanBench update analyzed the performance and efficiency of major Chinese-origin models, including Step 3.5 and Kimi-K2.5.

Details

The latest LisanBench update revealed an analysis of the performance and reasoning efficiency of major Chinese-origin models.

Step 3.5 Flash recorded a very impressive score. Although it uses a huge number of tokens, it showed excellent performance in terms of validity ratio and score. Kimi-K2.5 Thinking recorded a similar score to its predecessor K2, while its reasoning efficiency appears to have improved by roughly 2x.

In the Qwen3.5 series, the 397B A17B model outperformed GLM-5 and Minimax M2.5, but used a large number of tokens. Meanwhile, the 35B and 122B models were found to have bugs, such as the reasoning process being included in the output.

Key characteristics of other models are as follows:

  • Seed 2.0 Pro: Showed impressive performance in terms of reasoning efficiency.
  • GLM-5: Recorded lower performance than expected.
  • Minimax M2.5: Output format errors occurred, falling behind GPT-OSS-120B.
  • Seed 1.8, 2.0 Lite, Mini: Excellent price-to-performance ratio with very high token efficiency.

Currently, Opus 4.5 Thinking holds the No. 1 spot in the overall LisanBench rankings.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.