Grok 3 Beta — The Age of the Reasoning Agent
Key point
xAI has unveiled Grok 3 and Grok 3 mini, which maximize reasoning capability through reinforcement learning.
Details
xAI has released an early preview of Grok 3, which combines powerful reasoning capability with vast pretrained knowledge. It was trained on the Colossus supercluster using 10x the compute of the previous SOTA model, achieving a dramatic leap in math, coding, world knowledge, and instruction-following ability.
In particular, the Grok 3 (Think) model, refined through large-scale reinforcement learning (RL), spends anywhere from a few seconds to several minutes thinking through complex problems, correcting its own errors and exploring alternatives along the way. Users can use the Think button to view not just the model's final answer but its chain-of-thought reasoning process directly.
Key benchmark results are as follows:
- Chatbot Arena Elo score: 1402
- AIME 2025: 93.3% (with cons@64 applied)
- GPQA (graduate-level reasoning): 84.6%
- LiveCodeBench (code generation): 79.4%
xAI also introduced Grok 3 mini, built for cost-efficient reasoning. Optimized for STEM tasks, Grok 3 mini posted strong scores of 95.8% on AIME 2024 and 80.4% on LiveCodeBench.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.