Interaction models where humans and AI communicate naturally
Key point
ElevenLabs has unveiled models and features for real-time voice interaction.
Details
ElevenLabs aims for interaction models that converse with people in real time across audio, video, and text without turn-based delay. The flagship offerings are ElevenAgents and v3 Conversational, and it cited the Amanda case as an example, where an anxious customer on an urgent loan call is calmed and guided to resolution in under 2 minutes.
There are three conditions for success.
- Response at the sub-100ms level, with a target of sub-200ms for phone integration
- Turn-taking that considers both silence and spoken content together
- Expressive delivery that brings tone, pace, and emotion to life to fit the situation
What has already been released is concrete. Eleven v3 is the most expressive Text to Speech model, and Eleven v3 Conversational is the conversational version incorporated into ElevenAgents in February 2026; selecting it as the TTS model turns on turn-taking by default. Speculative turn-taking reduces perceived latency by starting LLM response generation early while the user is briefly silent.
On top of this, Flash v2.5 provides ultra-low-latency TTS with roughly 75ms inference, and Scribe v2 handles highly accurate Speech-to-Text. ElevenAgents Expressive Mode uses [laughs], [whispers], [sighs], [slow] tags to control delivery according to context. The figures are based on model inference time, and actual end-to-end latency varies depending on location and endpoint.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.