AI Briefing
KOSign in

ElevenLabs Scribe: Native Word-Level Timestamps and Event Tagging for Structured Audio Transcription

·2026.10.05 21:00

Key point

Scribe outputs structured token arrays with word, spacing, and audio_event types in a single API call, supporting up to 5 independent channels.

Details

ElevenLabs Scribe provides native word-level transcription that outputs structured arrays of tagged and timestamped token objects directly from audio input. Unlike conventional ASR models that produce text blocks requiring a secondary forced-alignment process for timing, Scribe calculates timestamps during transcription, reducing latency and compute overhead.

Token Structure and Event Tagging

The model categorizes every token into three types:

  • word: Recognized spoken words with start/end times and speaker IDs.
  • spacing: Explicitly tagged silence or pauses (excluded for languages without word spaces like Japanese or Mandarin).
  • audio_event: Non-verbal sounds such as laughter, applause, footsteps, or background music.

Enabling the tag_audio_events parameter allows the model to timestamp these non-speech events, enriching transcripts with context and enabling analytics on audience reactions or sentiment.

Real-Time and Multichannel Capabilities

Scribe supports real-time streaming via WebSocket with parameters like include_timestamps=true and include_language_detection=true. For multichannel audio, it can independently transcribe up to 5 channels in parallel. Speaker assignment is deterministic (Channel 0 maps to speaker_0), offering higher reliability than acoustic diarization for overlapping dialogue.

Practical Applications and Export Formats

Structured timestamped data enables precise keyword searching, compliance redaction via Entity Detection, and automated social media clipping. Scribe exports data in multiple formats:

  • JSON: For developers building interactive media or search databases.
  • SRT/VTT: For direct integration with video production tools like Adobe Premiere.
  • TXT/DOCX/PDF: For human-readable legal or audit documentation.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.