Independent Benchmarks Show Claude Opus 5.5 Leads in Intelligence but Trails in Security Efficiency
Key point
Opus 5.5 scores 58 on Artificial Analysis Intelligence Index but costs $5.98 per task compared to GPT-6 Astra's $3.26 due to higher token usage.
Details
Independent evaluations from Artificial Analysis and Endor Labs provide new performance metrics for Claude Opus 5.5, highlighting a trade-off between raw intelligence, cost efficiency, and security reliability.
Intelligence and Cost Efficiency
On the Artificial Analysis Intelligence Index, which aggregates ten tests, Opus 5.5 achieved a score of 58, surpassing GPT-6 Astra and Fable 5.1, which both scored 53. All models were evaluated at maximum effort settings.
- Terminal-Bench 4.0: Opus 5.5 and GPT-6 Astra performed nearly identically, with scores of 60% and 59% respectively.
- SciCode: Opus 5.5 demonstrated a significant advantage, leading Astra by 11 points.
- Cost per Task: Despite lower token prices, Opus 5.5 requires more tokens to complete tasks. At max effort, the cost per task is $5.98 for Opus 5.5, compared to $3.26 for GPT-6 Astra and $7.63 for Fable 5.1.
Security and Memorization Concerns
Endor Labs tested the models on real open-source projects involving code previously associated with security fixes, using hidden tests to verify safety without informing the model.
- Success Rates: Fable 5.1 achieved the highest secure completion rate at 37.4%, followed by Opus 5.5 at 33.5% and Opus 5 at 32.4%.
- Memorization Issue: Endor Labs observed that Opus 5.5 recalled known security fixes from its training data 51 times, more than any other model tested. These instances were excluded from the secure completion count. If included, Opus 5.5 would have ranked first in security by over 11 points.
- Performance Metrics: The full Endor Labs run cost $116 for Opus 5.5, significantly cheaper than $672 for Fable 5.1 and $1,116 for Opus 5. The median time per task was 2.2 minutes with no timeouts.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.