4.7 wins 69
Key point
In a 100-question blind evaluation, Opus 4.7 beat 4.6 by 69 to 30.
Details
Claude Opus 4.7 and Opus 4.6 were run under the same conditions via OpenRouter, then compared using 100 blind questions.
The evaluation was split into 5 categories — code, reasoning, analysis, communication, meta-alignment — and each answer was presented in randomized A/B order. Judging was handled by three models, GPT-5.4, Gemini 3.1 Pro, and DeepSeek V3.2, with the winner decided by majority vote.
The aggregate result was Opus 4.7 with 69 wins, 4.6 with 30 wins, and 1 tie. By judge, GPT-5.4 favored 4.7 at 69.7% and Gemini 3.1 Pro at 77.6%, but DeepSeek V3.2 more often picked 4.6, at 54 to 38.
4.7 was generally ahead in the per-category tallies as well.
- Code: 13 to 6, 1 tie
- Reasoning: 12 to 8
- Analysis: 16 to 4
- Communication: 14 to 6
- Meta-alignment: 13 to 7
The key takeaway is how much a single-judge leaderboard can shift. Even with the same questions and the same blind protocol, the conclusion changed depending on which model was used as the judge.
The conditions were also disclosed. Both models were accessed via OpenRouter, with generation settings set to temperature 0.7, max_tokens 4096. Judge settings used temperature 0.2, and 2 Gemini judgments were excluded due to structured output failures. The author noted that 100 questions are enough to get a sense of direction, but not enough to draw firm conclusions on detailed per-category results.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.