AI Briefing
KO

ElevenLabs-hosted LLMs

·2026.04.10 14:58

Key point

ElevenLabs has introduced self-hosted LLMs into its agent platform, reducing latency and improving cost efficiency for voice agents.

Details

ElevenLabs is introducing ElevenLabs-hosted LLMs into its agent platform to enable faster and more efficient voice agents. By hosting open source models directly on its own infrastructure, it achieves ultra-low latency, reduces inference costs, and allows customers to deploy agents without relying on separate external providers.

The key models are as follows:

  • GLM 4.5 Air: Delivers high reasoning accuracy and tool-calling performance, and can be operated at roughly 1/3 the cost of existing alternatives.
  • Qwen3-30b-a3b: Optimized for lightweight reasoning tasks, providing a natural conversational experience with a Time To First Sentence of under 150ms.

The core of this update is the co-located architecture. The hosted LLMs operate within the same environment as ElevenLabs' own STT (Speech to Text), TTS (Text to Speech), and turn-taking models. This integrated architecture reduces latency while strengthening reliability and data security.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.