Choose a recording
Add a supported audio or video file, or record directly from your microphone.
Turn spoken media into timestamped text, correct it in the editor, and download an SRT or WebVTT subtitle file.
Choose from the complete Whisper language list or let the tool detect the spoken language automatically. Support includes English, Spanish, Chinese, Arabic, Hindi, Japanese, Russian, Vietnamese, and many more.
This tool uses the open-source Whisper speech recognition model to transcribe audio and video locally. LocalScribe is not affiliated with OpenAI. Learn about Whisper transcription →
Your file stays private and is processed on your device.
Fast transcription is selected by default.
Add a supported audio or video file, or record directly from your microphone.
Select the spoken language or use automatic detection, then follow the progress on screen.
Correct the text and save it as TXT, SRT, or WebVTT subtitles.
Generate a timestamped transcript for interviews, lessons, presentations, podcasts, and videos, then review it before export.
Download a widely supported subtitle file with numbered, timestamped caption blocks.
Export WebVTT for websites, web video players, and supported publishing platforms.
Review names, numbers, and specialist terms before downloading your final text.
Choose from 99 supported transcription languages for international recordings.
Your selected recording and completed transcript stay on your device. The transcription tool does not upload them to our servers.
Select the video above, run transcription, review the text, and choose SRT from the download buttons.
Both store caption text and timestamps. SRT is widely supported by editing and publishing tools, while WebVTT is designed for web video.
You can edit the transcript text in the page before export. Fine-grained timestamp editing is not yet available.
No. It creates separate SRT and VTT subtitle files; it does not render text into the video image.