A new version isn't always better
Key point
Unlike general-purpose benchmarks, Opus 4.7 scored lower than 4.6 in a real-world workflow.
Details
Comparing claude-opus-4.6 and claude-opus-4.7 on a reasoning/logic test for a real SaaS workflow, 4.7 — despite being the newer version — lagged behind in accuracy.
- claude-opus-4.6: 66% (47.0/71.0), cost $0.0257, time 44.50s
- claude-opus-4.7: 61% (43.0/71.0), cost $0.0170, time 36.56s
4.7 was cheaper and faster, but for this user's specific agentic flow, its accuracy and stability were lower than 4.6's. This highlights that a model that looks good on general-purpose evaluations isn't guaranteed to be better on real work tasks.
The author also says they've noticed a performance regression in Opus 4.6 over the past few months, and that eval data saved early in the release let them confirm that regression. The conclusion: task-model matching matters more than which model is "better."
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.