AI Briefing
KO

The 3-Second Voice Revolution

·2026.04.10 14:58

Key point

Speech synthesis has rapidly evolved from the Hawking voice to 3-second voice cloning.

Details

The flow of speech synthesis is laid out at a glance.

  • Dennis Klatt's voice was used as Stephen Hawking's voice, becoming a symbol of early TTS.
  • WaveNet pushed quality forward, recording 4.21 vs 3.86 compared to the previous best system.
  • Tacotron 2 scored 4.53, coming close to real human speech at 4.58.
  • VALL-E could clone a speaker's voice from just a 3-second sample, after which Microsoft refused to release VALL-E 2.
  • Kokoro is mentioned as being able to compete with ElevenLabs with 82M parameters and a $400 training cost.
  • In a 2025 Nature study, people rated AI voices as more trustworthy than real voices.

The core message is clear. Voice technology has moved beyond simple synthesis into a stage that handles short-sample-based cloning and even psychological trust.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.