GoBench Released, Gaining Attention as an LLM Reasoning Benchmark
Key point
GoBench, a benchmark evaluating LLM Go skills, has been released and is gaining attention as a reasoning metric due to its high correlation with ARC-AGI 2.
Details
A new benchmark GoBench has been released to evaluate the general reasoning capabilities of LLMs. This benchmark evaluates models by playing 9x9 Go games against KataGo at various levels (from random to superhuman).
Key features and results are as follows:
- High correlation: GoBench scores show a strong correlation of r=0.83 with the ARC-AGI 2 benchmark, which measures complex reasoning abilities.
- Unsaturated state: Unlike existing benchmarks that have reached saturation, GoBench maintains high discriminative power.
- Model performance: GPT-6 Astra max recorded 2500 Elo, which is significantly lower than the 4400 Elo of top-tier KataGo. However, Codex with Astra, which allowed coding tools and 2 hours of preparation, reached 3560 Elo.
The developers plan to continuously update the leaderboard as long as the benchmark does not become saturated.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.