AI Briefing
KO

Challenges and Opportunities of Synthetic Voice

·2024.03.29 09:00

Key point

OpenAI has unveiled early use cases and a safe deployment approach for **Voice Engine**, which generates natural-sounding voices from just a 15-second sample.

Details

Voice Engine is a model that generates natural, emotive speech highly similar to the original speaker using only text input and just 15 seconds of audio sample. OpenAI has used this technology in the text-to-speech API, ChatGPT Voice, Read Aloud, and more, and is currently conducting a small, cautious preview while exploring ways for society to adapt, given the potential for misuse.

OpenAI is currently testing the technology's potential across various industries with trusted partners. Key use cases include:

  • Education: Age of Learning supports children's learning through personalized voice content and real-time interaction.
  • Translation: HeyGen translates videos into multiple languages while preserving the speaker's unique accent.
  • Global services: Dimagi supports local languages such as Swahili to help empower healthcare workers in remote areas.
  • Accessibility and Healthcare: Livox provides customized voices for non-verbal users, and Lifespan is researching voice restoration for patients who have lost their voice due to illness.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.