A detailed look at how Voice Engine works and its safety research
Key point
OpenAI has revealed the technical principles behind Voice Engine, which generates custom voices from a 15-second audio sample, along with details of its safety research.
Details
Voice Engine is a TTS (Text-to-Speech) model that generates human-like audio from just text and a 15-second audio sample. This model does not undergo any separate fine-tuning or model customization process for a specific speaker.
Instead, it leverages a Diffusion process. Starting from random noise, it generates audio by progressively removing noise, allowing it to precisely reproduce the intonation, accent, and speaking style of the speaker captured in the 15-second sample.
Development of this model began in late 2022, and it has gone through internal testing and technical demonstrations for policymakers. It is currently applied in limited form to ChatGPT Voice Mode and the TTS API, and has recently been made available in preview form to trusted partners for custom voice generation.
OpenAI is pursuing the following safety goals to manage the risks of synthetic voice technology:
- Phasing out voice-based authentication for security purposes, such as accessing bank accounts
- Exploring policies to protect individuals' voice rights in the AI era
- Raising public understanding of deceptive AI content
- Developing and accelerating the adoption of technology to trace the provenance of audio content
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.