AI Briefing
Sign in

Anthropic Releases Claude Sonnet 5.5 with Major Agentic Coding and Cost Efficiency Gains

·2026.09.28 09:00

Key point

Claude Sonnet 5.5 achieves 70.6% on Terminal-Bench 4.0, a massive jump from Sonnet 5's 10.3%, while maintaining the same pricing as its predecessor.

1 / 3

Details

Anthropic has released Claude Sonnet 5.5, the second model in the Claude 5.5 series, positioning it as a faster and more cost-effective complement to Opus 5.5. While Opus 5.5 remains superior for complex, open-ended judgment tasks, Sonnet 5.5 is optimized for well-defined daily work, bug fixes, and document generation. A high-volume, cost-sensitive variant, Claude Haiku 5.5, is expected to launch within weeks.

Performance Benchmarks

Sonnet 5.5 demonstrates significant improvements over Sonnet 5 across key benchmarks, often approaching Opus 5.5 performance at a lower cost:

  • Terminal-Bench 4.0 (Agentic Coding): Scored 70.6%, a substantial increase from Sonnet 5's 10.3% (Opus 5.5 scored 66.4%).
  • GDPval-AA v2.1 (Knowledge Work): Achieved 1844 points, nearly matching Opus 5.5's 1846 points and surpassing Sonnet 5's 1449 points.
  • OSWorld 2.1 (Computer Use): Reached 80.1%, up from Sonnet 5's 57.0% (Opus 5.5: 81.8%).
  • Chartography (Visual Chart Recognition): Scored 61.6%, a dramatic rise from Sonnet 5's 15.6%.
  • FrontierCode 1.1: Recorded 52.1% at Xhigh effort, compared to Sonnet 5's 42.4%.
  • Pokémon Red: Became the first Sonnet model to clear the game using only screenshots.

Cost and Speed Efficiency

Pricing remains identical to Sonnet 5 ($2/1M input tokens, $10/1M output tokens), but efficiency gains reduce task costs by up to 30% due to lower token usage. Output generation is 30%+ faster than Sonnet 5, making it the fastest Sonnet model to date. At Low/Medium effort settings, it achieves Sonnet 5's peak scores at approximately 1/10th the cost per task.

Safety and Guardrails

Sonnet 5.5 introduces enhanced safety measures, particularly in cybersecurity and reasoning extraction:

  • Cybersecurity: Features Opus 5-level capabilities with guardrails similar to Opus 5.5. High-risk cyber tasks fall back to Sonnet 5, while advanced access is available via the Cyber Verification Program.
  • Distillation Protection: It is the first Sonnet model with safety classifiers to prevent reasoning extraction, ensuring thinking processes remain separated from accounts.
  • Biology: Maintains Sonnet 5-level guardrails, blocking only high-risk requests while allowing most research and clinical work.

Enterprise Feedback

Early testers reported significant operational improvements:

  • Early testers: Achieved Opus 5-equivalent scores in 118 app builds with fewer iterations (3.6 vs. 7.7 for Opus 5).
  • Early testers: Ticket processing speed improved by 20% with fewer incorrect decisions.
  • Balyasny Asset Management: Used ~121k tokens per answer vs. 497k for Sonnet 5 in financial tasks.
  • Early testers: Reported 2.4x faster speeds and a 12% reduction in total token usage.

The model is available on AWS, Google Cloud, and Microsoft Azure under the model name claude-sonnet-5-5, with Zero Data Retention options supported.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.