All tools

MP4 to text

Extract the spoken words from an MP4 video as text, ready to copy or export as subtitles.

Transcribed on your device

Local transcription benchNo file loaded
TranscribeWhisper, on-device

Speed depends on your device. The model downloads once, then works offline.

How it works

  1. Choose an MP4 video — the audio track is extracted automatically.
  2. Pick the spoken language and a model, then transcribe.
  3. Read the transcript as it appears, then copy it or download TXT, SRT, or VTT.

Getting text out of an MP4

MP4 is where spoken content accumulates: meeting recordings, lecture captures, interviews, screen recordings, downloaded webinars. Watching them back at 1x to take notes is the slow way — a transcript is searchable, skimmable, and quotable. This tool reads the audio track straight out of the MP4 container (no separate conversion to MP3 needed) and transcribes it with an open Whisper model that runs entirely in your browser tab.

Because nothing is uploaded, there is no file-size cap and no queue: a two-hour meeting recording is limited only by your device, not by a free tier. The first transcription downloads a small model (~40 MB) that is cached for every later visit — after that, the tool even works offline. Speed comes from your hardware: a recent laptop processes several times faster than real-time, so a one-hour video takes a few minutes, with text appearing window-by-window as it is recognized.

What you get out

For the best result, pick the spoken language before starting and use the sharper model on noisy recordings. If you only need part of the video, trim it first with the video trimmer — less audio means a faster, cleaner transcript. The transcription guide covers accuracy expectations in more detail.

Questions

Do I need to convert the MP4 to MP3 first?
No. The audio track is extracted from the MP4 container directly in your browser, then transcribed. Dropping in the video file itself is the intended workflow.
Is there a length or file-size limit?
No fixed limit — the video is processed in 30-second windows on your own machine, so long recordings simply take proportionally longer. You can stop at any point and keep the transcript produced so far.
Can I get subtitles from my MP4?
Yes. After transcribing, switch the result to the SRT or VTT tab and copy or download the file — the timestamps are included, ready for editors and players. See the video to SRT page for the subtitle-focused workflow.
Is the video uploaded anywhere?
No. The audio is extracted and transcribed inside your browser tab using a locally running model. The only network request ever made is the one-time model download.
What about MP4s with music or multiple speakers?
Background music and crosstalk reduce accuracy for any small on-device model — use the sharper model and expect to proofread. Speakers are not separated or labeled; the transcript is one continuous text.