MODELS
ElevenLabs Scribe v2
by ElevenLabs
modelsource:elevenlabs.iospeech-to-texttranscriptiondiarizationtimestampsaudio-taggingElevenLabs
Overview
Speech-to-text model for transcription across 90+ languages with word-level timestamps, speaker diarization, keyterm prompting, and audio tagging.
Details
ElevenLabs Scribe v2 is a speech-to-text model from ElevenLabs. The supplied ElevenLabs documentation and help sources describe it as supporting transcription across 90+ languages, with word-level timestamps, speaker diarization, keyterm prompting, and audio tagging. It is documented in ElevenLabs model documentation and Speech-to-Text help materials.
When to Use
Use when you need speech-to-text transcription across 90+ languages. Use when transcription output needs word-level timestamps speaker diarization keyterm prompting or audio tagging.
Getting Started
- Open the ElevenLabs model documentation at https://elevenlabs.io/docs/models.
- Review the ElevenLabs Speech-to-Text help material for Scribe v2 capabilities.
- Test Scribe v2 on representative audio before using it in production workflows.
Key Features
- •Transcription across 90+ languages
- •Word-level timestamps
- •Speaker diarization
- •Keyterm prompting
- •Audio tagging
Capabilities
- •speech-to-text
- •transcription
- •speaker diarization
- •word-level timestamps
- •audio tagging
Last updated Jun 9, 2026