Opus 4.7 Long-Context Regression
·2026.04.17 00:46
Key point
Opus 4.7 improved on coding but regressed significantly on long-context reasoning.
Details
In a system card comparison, Opus 4.7 improved on the SWE-bench series, but Opus 4.6 was superior on the remaining 5 items.
- SWE-bench Verified: 80.8% → 87.6%
- SWE-bench Pro: 53.4% → 64.3%
On the other hand, the search and long-context reasoning series declined.
- BrowseComp (10M token): 83.7% → 79.3%
- DeepSearchQA F1: 91.3% → 89.1%
- MRCR v2 8-needle @ 256k: 91.9% → 59.2%
- MRCR v2 8-needle @ 1M: 78.3% → 32.2%
- ARC-AGI-1: 93.0% → 92.0%
The drop is especially large in MRCR v2 @ 1M, which reads as a signal that 4.7's ability to accurately find specific information in long-context has weakened significantly.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.