Choose a recording
Add a supported audio or video file, or record directly from your microphone.
Separate different voices as Speaker 1, Speaker 2, and more while your recording remains on your device.
Choose from the complete Whisper language list or let the tool detect the spoken language automatically. Support includes English, Spanish, Chinese, Arabic, Hindi, Japanese, Russian, Vietnamese, and many more.
This tool uses the open-source Whisper speech recognition model to transcribe audio and video locally. LocalScribe is not affiliated with OpenAI. Learn about Whisper transcription →
Your file stays private and is processed on your device.
Fast transcription is selected by default.
Add a supported audio or video file, or record directly from your microphone.
Select the spoken language or use automatic detection, then follow the progress on screen.
Correct the text and save it as TXT, SRT, or WebVTT subtitles.
Optional speaker diarization analyzes voice characteristics and aligns anonymous labels with the timestamped transcript. It does not recognize names or identities.
Distinguish participants in recorded calls and discussions when reviewing notes and decisions.
Separate questions and answers into an easier-to-read interview transcript.
Label changing voices in episodes, roundtables, and recorded conversations.
Download speaker-prefixed text and timestamped subtitle cues after reviewing the result.
Your selected recording and completed transcript stay on your device. The transcription tool does not upload them to our servers.
Speaker diarization estimates who spoke when. This tool uses anonymous labels such as Speaker 1 and Speaker 2; it does not identify people by name.
Yes. When enabled, both transcription and speaker analysis run on your device. Additional speaker models are downloaded and cached by the browser the first time.
Quality depends on the recording. Overlapping speech, background noise, short replies, and similar voices can cause mistakes, so review labels before publishing.
The optional feature performs separate voice segmentation and comparison after transcription. Its extra models and processing are only used when you enable speaker labels.