Introducing gpt-realtime and updates to the Realtime API
Key point
OpenAI has unveiled gpt-realtime, an enhanced voice conversation model, along with an upgraded Realtime API with more features.
Details
OpenAI has officially launched (GA) the Realtime API, introducing the more advanced voice conversation model gpt-realtime. This update allows developers to build more reliable, production-grade voice agents.
The new API includes support for MCP (Model Context Protocol) servers, image input, and phone connectivity via SIP (Session Initiation Protocol). This enables voice agents to access additional tools and context, allowing for even more powerful capabilities.
Unlike previous approaches that chained together separate STT (speech-to-text) and TTS (text-to-speech) models, gpt-realtime uses a single model to directly process and generate audio. This approach reduces latency and preserves the subtle nuances of speech, enabling much more natural conversations.
Key improvements include:
- Audio Quality: More natural intonation and emotional expression are now possible, with two new voices, Cedar and Marin, added.
- Intelligence and Understanding: The model can pick up on non-verbal cues such as laughter and switch languages mid-sentence. It achieved 82.8% accuracy on the Big Bench Audio benchmark, significantly outperforming the previous model (65.6%).
- Instruction Following and Tool Use: Complex command execution and precise Function Calling capabilities have been greatly enhanced.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.