AI Gateway Begins Supporting Real-Time Voice, Speech, and Transcription Features
Key point
Vercel AI Gateway now supports real-time voice agents, TTS, and STT capabilities, available in beta through AI SDK 7.
Details
Vercel's AI Gateway has begun supporting voice and audio models. Developers can now build real-time voice agents, convert text to speech (TTS), and transcribe audio to text (STT).
Through this update, the same Observability, Spend controls, and BYOK (Bring-Your-Own-Key) features available for existing text, image, and video models are now provided, with no additional markup or platform fees. This feature is currently in beta and available through AI SDK 7.
Key features include the following:
- Realtime voice agents: Through low-latency conversation, they listen to users and respond instantly, and can call Tools mid-conversation to perform tasks.
- Text to Speech (TTS): Generates speech from text using a selected voice and output format such as MP3.
- Speech to Text (STT): Transcribes audio recordings into text via file buffers, base64 strings, or URLs.
Developers can easily manage microphone capture and audio playback using the useRealtime hook in the AI SDK, and can test features by talking directly with models in the AI Gateway Playground without writing any code.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.