AI Briefing
KO

OpenAI Unveils Next-Generation Audio Models for the API

·2025.03.20 20:00

Key point

OpenAI has launched next-generation audio models via the API, capable of sophisticated speech recognition and customized speech synthesis.

Details

OpenAI has released new speech-to-text (STT) and text-to-speech (TTS) audio models through the API to support more intuitive interactions. This allows developers to build more powerful and customized voice agents.

The new STT models, gpt-4o-transcribe and gpt-4o-mini-transcribe, recorded improved WER (Word Error Rate) compared to the existing Whisper model. They show high accuracy and reliability even in challenging conditions such as accents, noisy environments, and varying speech speeds.

For the TTS model, the ability for developers to instruct specific speaking styles has been introduced for the first time. For example, instructions like "speak like an empathetic customer service agent" enable emotionally expressive speech, allowing for a wide range of applications from creative storytelling to customer service.

Key features include the following:

  • Introduction of gpt-4o-transcribe and gpt-4o-mini-transcribe models
  • Demonstrated performance surpassing Whisper v2/v3 on the FLEURS benchmark
  • Improved accuracy and language recognition capability through reinforcement learning and high-quality datasets
  • Ability to generate customized voices reflecting specific characters or emotions

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.