ElevenLabs
Scribe v2
ElevenLabs' most accurate transcription: word timings, speaker labels and 90+ languages, from a sound or a video. Billed per started minute of the recording. On Leap, it sits beside 131 other models in one studio.
- Speech to text
- Captions
- Start free, no card
- Every top model, one account
- Pay per run, priced before it starts
- 01Describe what you want to see
- 02See the exact price, then press go
- 03Keep it in your library, or take it further
Start with an idea
Bring a photo to Scribe v2
Scribe v2 works on a picture you already have. Open it in the studio, add your photo, and see the result in seconds.
Questions
Scribe v2 on Leap, answered.
- Is this the real Scribe v2?
- Yes. Leap sends your request to ElevenLabs's model through its provider and gives you back what it makes. Switch to any other model in the same studio when an idea needs it.
- What does it cost?
- In dollars, per run: the provider's list price plus 10%. You see the exact price before every run, a run that fails costs nothing, and there is no subscription. Every model's price is on the pricing page.
- Do I need a ElevenLabs account?
- No. One Leap account and one balance run Scribe v2 and every other model, in the studio, through the API and from your agents.
- What happens to what I make?
- It stays in your library while your account is active, ready to use anywhere.
Your idea, Scribe v2.
Start free, no card. Every run shows its price before it starts.
Make it with Scribe v2For developers
Run Scribe v2 from code
- Model ID
- elevenlabs/scribe-v2
- Makes
- Speech to text
- Pricing
- Per run, quoted first: POST /v1/quotes
Inputs
What the model takes. Leave out anything with a default.
Recordingaudio_file
An imageThe sound to transcribe: an uploaded file, by ID. Up to 4 hours.
Videovideo
An imageOr a video whose speech to transcribe: an uploaded file, by ID. Up to 4 hours.
Languagelanguage
Detect, English, Spanish, Portuguese, French, German, Italian, Dutch, Polish, Russian, Ukrainian, Turkish, Arabic, Hebrew, Persian, Hindi, Bengali, Urdu, Tamil, Indonesian, Malay, Filipino, Vietnamese, Thai, Chinese, Japanese, Korean, Swedish, Danish, Norwegian, Finnish, Greek, Czech, Romanian, Hungarian, Swahili, Yoruba. Default: DetectThe language spoken. Detect works for most recordings.
Speaker labelsspeakers
On or off. Default: onTell the voices apart: Speaker 1, Speaker 2.
Sound tagssettings.sound_events
On or off. Default: offMark laughter, applause and music in the transcript.
The request
The same envelope for every model; only the model and its input change. With prefer: wait=60 the answer holds the result. Agents can call it through the MCP server at mcp.tryleap.ai.
curl https://api.tryleap.ai/v1/generations \ -H "x-api-key: $LEAP_API_KEY" \ -H "content-type: application/json" \ -H "prefer: wait=60" \ -d '{ "model": "elevenlabs/scribe-v2", "input": {} }'