AI Briefing
KO

OpenAI Unveils 3 Voice Models

·2026.05.08 21:50

Key point

OpenAI added 3 voice models to the Realtime API.

Details

OpenAI added 3 audio models to the Realtime API.

  • GPT-Realtime-2: Built on GPT-5-level reasoning, it supports long conversations, interruption recovery, tool calling, and responses while tasks are in progress.
  • GPT-Realtime-Translate: Translates 70+ input languages into 13 output languages in real time.
  • GPT-Realtime-Whisper: Handles streaming speech transcription.

OpenAI is pushing voice not as an add-on feature for chatbots but as the interface itself. It emphasized workflows that connect with apps to perform actions while a conversation continues, such as customer support, scheduling, travel guidance, and meeting translation.

Examples cited included Zillow's home search and tour booking, Deutsche Telekom's multilingual customer support, Priceline's conversational travel planning, and Vimeo's real-time broadcast translation.

Pricing for Realtime-2 is $32 per 1 million input audio tokens and $64 per 1 million output audio tokens. Translate is $0.034 per minute, and Whisper is $0.017 per minute, and they can be tested right away in the Playground.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.