AI Briefing
KOSign in

ElevenLabs Guides Meeting Transcription Setup with Scribe v2 and Scribe v2 Realtime

·2026.10.08 21:00

Key point

Scribe v2 Realtime achieves 150 ms latency for live captioning, while batch Scribe v2 supports up to 1,000 keyword terms for accurate post-meeting transcripts.

Details

ElevenLabs provides a guide for building meeting transcription products using its Scribe v2 and Scribe v2 Realtime models. These APIs convert audio into time-stamped, speaker-labeled text, addressing the limitations of manual note-taking and generic commercial tools. The solution supports 90+ languages and offers features like Zero Retention Mode for enterprise privacy compliance.

Real-Time vs. Batch Transcription

The article distinguishes between two primary workflows based on latency and accuracy needs:

  • Real-time transcription: Uses Scribe v2 Realtime for live captioning and in-meeting bots. It achieves low latency of 150 ms between audio input and text output. This method requires an active WebSocket connection and supports up to 50 key terms for prompting.
  • Batch transcription: Uses Scribe v2 for post-meeting notes and record-keeping. It processes entire recordings asynchronously via REST API, offering higher accuracy due to full conversational context analysis. This approach supports up to 1,000 keyword terms for improved recognition of specific industry or brand vocabulary.

Implementation Details

For real-time integration, developers open a WebSocket to stream audio and receive partial or final transcripts. The API supports two streaming methods: client-side streaming using single-use tokens to protect API keys, and server-side streaming using the main API key. Commit strategies include manual commits (recommended every 20–30 seconds) or Voice Activity Detection (VAD) to automatically segment speech based on silence.

Batch transcription involves uploading audio files to the Scribe v2 REST endpoint and configuring a webhook to receive results. This method supports multichannel transcription, which tags individual audio channels with unique speaker identifiers, making it suitable for conferences, court recordings, and podcasts. The model also handles non-speech events like laughter and footsteps through dynamic audio tagging.

Addressing Common Challenges

The API mitigates typical transcription issues through specific features:

  • Overlapping conversations: Handled via diarization and multichannel transcription to separate speakers.
  • Misinterpretations: Corrected using keyword prompting, which guides the model to recognize specific terms like "DevOps" or brand names in context.
  • Multilingual speakers: Supported by the model's ability to transcribe across 90+ languages, accommodating code-switching and mixed-language meetings.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.