AI Briefing
Sign in

OpenAI Prepares Wider Rollout of Ultrafast API Mode Ahead of DevDay

·2026.09.26 21:24

Key point

The Ultrafast mode, powered by Cerebras, delivers up to 750 output tokens per second and is currently limited to selected customers.

1 / 3

Details

OpenAI is preparing to expand access to its Ultrafast API mode, with new references appearing in platform documentation ahead of DevDay on September 29. TestingCatalog identified a hidden speed selector in the Responses API Playground that allows developers to choose between Standard, Fast, and Ultrafast processing options.

Performance and Infrastructure

The Ultrafast mode was officially previewed with GPT-5.6 Sol, offering speeds of up to 750 output tokens per second and inference up to 14× faster than Standard. OpenAI confirmed that this mode is powered by Cerebras, aligning with a previously announced 750 MW partnership where capacity is being deployed in stages through 2028.

Future Model Support and Economics

While the feature is currently restricted to selected customers, the recent release of the GPT-6 family (including GPT-6 Sol and GPT-6 Astra) raises questions about broader model support, though no confirmation exists that every GPT-6 model will support Ultrafast at launch. For developers, the tiered system offers a trade-off between cost and latency, allowing organizations to reserve higher-cost inference for workloads where response time directly impacts revenue or productivity.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.