GPT-5.5 Released
Key point
GPT-5.5 has been released as a new model that handles coding, knowledge work, and scientific research faster and more efficiently.
Details
GPT-5.5 has been released as a practical, work-ready model that can be entrusted with coding, online research, data analysis, document and spreadsheet creation, and even software operation. It is designed to plan on its own, use tools, and follow through with verification and revision even when handed complex, multi-step tasks.
Improvements are especially notable in agentic coding, computer use, knowledge work, and early-stage scientific research. It maintains per-token latency similar to GPT-5.4 while delivering higher intelligence, and on Codex tasks it completes the same work with fewer tokens, improving efficiency as well. On Artificial Analysis's Coding Index, it reportedly delivers top-tier performance at roughly half the cost of competing frontier coding models.
For safe deployment, the strongest guardrails to date have been applied. The release was prepared through internal and external red team testing, additional testing on cybersecurity and biology, and feedback from around 200 early access partners. The rollout has begun in ChatGPT and Codex for Plus, Pro, Business, and Enterprise, with GPT-5.5 Pro rolling out first to ChatGPT's Pro, Business, and Enterprise tiers. The API version will be available soon after meeting separate safety requirements.
- Coding: Terminal-Bench 2.0 82.7%, SWE-Bench Pro 58.6%, Expert-SWE 73.1%, CyberGym 81.8%
- Knowledge/work: GDPval 84.9% win-or-tie rate, OSWorld-Verified 78.7%, Tau2-bench Telecom 98.0% (without prompt tuning), FinanceAgent 60.0%, internal investment banking modeling task 88.5%, OfficeQA Pro 54.1%
- Additional comparisons: BrowseComp 84.4% (Pro 90.1%), FrontierMath Tier 1–3 51.7% (Pro 52.4%), FrontierMath Tier 4 35.4% (Pro 39.6%)
Real-world use cases are also concrete. Over 85% of people inside OpenAI use Codex weekly, and it has been applied to automating request classification for the communications team, reviewing 24,771 K-1 tax documents (71,637 pages) in Finance, and automating weekly reports for the GTM team. GPT-5.5 Thinking provides faster, more concise answers to high-difficulty questions, and GPT-5.5 Pro delivers more comprehensive and accurate results on business, legal, education, and data science tasks.
In scientific research, it showed improvements on GeneBench and BixBench, and in internal harnesses it also contributed to finding new proofs for Ramsey numbers.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.