MP4 to text
Extract the spoken words from an MP4 video as text, ready to copy or export as subtitles.
Transcribed on your device
How it works
- Choose an MP4 video — the audio track is extracted automatically.
- Pick the spoken language and a model, then transcribe.
- Read the transcript as it appears, then copy it or download TXT, SRT, or VTT.
Getting text out of an MP4
MP4 is where spoken content accumulates: meeting recordings, lecture captures, interviews, screen recordings, downloaded webinars. Watching them back at 1x to take notes is the slow way — a transcript is searchable, skimmable, and quotable. This tool reads the audio track straight out of the MP4 container (no separate conversion to MP3 needed) and transcribes it with an open Whisper model that runs entirely in your browser tab.
Because nothing is uploaded, there is no file-size cap and no queue: a two-hour meeting recording is limited only by your device, not by a free tier. The first transcription downloads a small model (~40 MB) that is cached for every later visit — after that, the tool even works offline. Speed comes from your hardware: a recent laptop processes several times faster than real-time, so a one-hour video takes a few minutes, with text appearing window-by-window as it is recognized.
What you get out
- A timestamped transcript you can read while the rest is still processing — stop early and keep the partial text.
- Three export shapes — TXT for notes and documents, SRT for video editors, VTT for embedding subtitles on the web. Switching between them is instant.
- One-click copy of the whole transcript in whichever format you have open.
For the best result, pick the spoken language before starting and use the sharper model on noisy recordings. If you only need part of the video, trim it first with the video trimmer — less audio means a faster, cleaner transcript. The transcription guide covers accuracy expectations in more detail.
Questions
- Do I need to convert the MP4 to MP3 first?
- No. The audio track is extracted from the MP4 container directly in your browser, then transcribed. Dropping in the video file itself is the intended workflow.
- Is there a length or file-size limit?
- No fixed limit — the video is processed in 30-second windows on your own machine, so long recordings simply take proportionally longer. You can stop at any point and keep the transcript produced so far.
- Can I get subtitles from my MP4?
- Yes. After transcribing, switch the result to the SRT or VTT tab and copy or download the file — the timestamps are included, ready for editors and players. See the video to SRT page for the subtitle-focused workflow.
- Is the video uploaded anywhere?
- No. The audio is extracted and transcribed inside your browser tab using a locally running model. The only network request ever made is the one-time model download.
- What about MP4s with music or multiple speakers?
- Background music and crosstalk reduce accuracy for any small on-device model — use the sharper model and expect to proofread. Speakers are not separated or labeled; the transcript is one continuous text.