Overview
Transcription converts the spoken audio in your video into timestamped captions in the source language. Powered by ElevenLabs Scribe v2, it produces word-level timing and speaker diarization — the essential foundation for translation and dubbing.Starting a transcription
- Open your video from the dashboard
- Click Transcribe
- Select the spoken language, or choose Auto-detect
- Click Start
Supported languages
Neolli supports transcription in 10 languages (plus auto-detect):
For the full capabilities matrix across all features, see Supported Languages.
Features
- Speaker diarization — Automatically identifies and labels different speakers
- Word-level timing — Each word gets its own precise timestamp for accurate syncing
- Auto-detect — Identifies the spoken language automatically for common languages
Processing time
You can close the browser while a job is running — it continues in the background. The dashboard shows job progress in real time.
File requirements
- Max file size: 3 GB
- Supported formats: MP4, MOV, MKV, AVI, WebM, and most common video/audio formats
- Audio: Must contain a detectable audio track with speech
After transcription
Once complete, you can:- Review and edit captions in the caption editor
- Add target languages for translation or dubbing
- Export the source captions as SRT