How to transcribe a video file to text for free (with speaker labels)

An hour-long meeting recorded as an mp4 takes roughly 4 hours to transcribe by hand - and that is before you try to work out who said what. Most tools that accept video either strip the audio and charge per minute, or lock transcription behind a subscription. WhoSayAh transcribes video files directly in your browser: no conversion, no account, no charge for files within the size and length limits, and automatic speaker labels so every line is attributed to the person who said it.

What video formats are accepted

Common video containers upload as-is. Only the audio track is read during transcription, so a 500 MB screen recording of a 30-minute call produces the same result as an audio-only export of the same call - without any extra conversion step.

CategoryAccepted formatsTypical size for 30 minutes
Videomp4, mov, mkv, webm200-600 MB
Audiomp3, m4a, wav, ogg, opus, flac, aac20-60 MB

The upload limit is 100 MB and 120 minutes per file. Both are checked immediately at upload time, so an over-sized file is rejected in seconds rather than after a long wait. If you have a longer recording - a full conference session, a deposition, a day of interviews - split it into segments under the limit. Most free video editors can cut by time marker without re-encoding.

Video files frequently exceed the size limit even for moderate durations, because a screen recording at high resolution can run to several hundred megabytes for a single hour. The quickest fix is to extract the audio track before uploading. On a Mac, QuickTime Player exports audio only via File → Export → Audio Only. On Windows and Linux, the free VLC media player converts a video to .m4a in under a minute via Media → Convert/Save, selecting an audio-only profile. The transcript is identical; the file is a fraction of the size.

Step by step: from video file to labelled transcript

  1. Open the home page and drop your file. A dialog appears showing the file name, size and detected length. If the file is over 100 MB or 120 minutes, the dialog says so immediately.
  2. Confirm the language. Leave Detect automatically on for most recordings. Select the language explicitly when the recording is under 2 minutes or the environment was noisy - auto-detection has less to go on in those cases.
  3. Set the speaker count. Keep Label speakers on and enter how many people actually spoke - not how many attended. A 12-person all-hands where 4 people contribute is a 4-speaker recording. If you are not sure, press I don't know - detect it.
  4. Press Transcribe and read along. Lines appear on screen as they are processed. A 20-minute recording is typically readable within 2 to 3 minutes; a full hour within 6 to 8 minutes. You do not need to stay on the page, but results appear there.
  5. Name the speakers. When the transcript is complete, press Name speakers. Each voice is listed with the total time it spoke and its first line. For a meeting you just attended, placing everyone takes about 30 seconds. The names replace Speaker 1, Speaker 2 across the entire transcript in one step.
  6. Download or copy. Take the transcript as .txt (plain text, one line per speaking turn, name prefixed) or .vtt (WebVTT with timestamps, ready for video software or YouTube). Or copy the whole thing to clipboard in one click.

Why you do not need to convert the video first

Many transcription tools accept video in name only: they re-encode the file server-side before processing it, adding a format-conversion step you do not see. WhoSayAh reads the audio track directly from the container, skipping that conversion. In practice:

The one case where pre-conversion still makes sense is when the video file is very large. A 2-hour screen recording at 1080p is often 1.5-3 GB, which will exceed 100 MB. There, the right move is to extract the audio - not to compress the video. The audio track from that same recording is typically 50-120 MB.

Speaker labels: how they work and when they help most

A transcript without speaker labels is a wall of text. You can read every word and still not know who made the commitment, who raised the objection, or who owns the follow-up. Speaker labels assign each line to the person who said it, automatically, without you listening back and annotating by hand. For a 1-hour recording, that saves 3 to 4 hours of work.

Speaker labels work from the audio content of the file, not from metadata or separate microphone tracks. This means they work with any video recording - a Zoom mp4, a Teams meeting export, a phone video of an in-person session. Attribution is strongest when:

Once the transcript is complete, Name speakers lists each voice with how long it spoke and its first line - enough to identify everyone in a meeting you just attended. The names go into both the .txt and .vtt downloads. For a two-person interview with a remote guest, the interview transcription guide covers the specific settings that matter most for that format, including how to handle volume imbalance between an interviewer and a remote guest.

Output formats: .txt, .vtt and copy

Once the transcript is ready there are three ways to take the text:

WebVTT is a web standard, so the .vtt file works across platforms without any additional tooling or conversion. Teams that publish recorded video at scale - a legal firm archiving depositions, a media company captioning interviews, an HR department releasing training sessions - reach a point where uploading one file at a time is not practical. CrocOTT handles the VOD-platform side of that: captions synced across a full catalogue, accessibility compliance built in, no manual upload loop.

What happens to your video file

The uploaded file is deleted as soon as the transcript is ready. The transcript is deleted after 24 hours. There is no account, so nothing links either to you. The privacy policy states exactly what is stored and for how long. The terms of use carry the upload size and length limits in full.

For recordings that cannot leave your own network - NDA-covered client calls, medical consultations, legal depositions, anything subject to a data-residency regulation - the alternative is a self-hosted installation where the audio never reaches an external server. FastoCloud provides the infrastructure for that: the same browser-based upload experience described in this guide, running entirely inside your own environment with no per-minute charge.

Common questions

Does the video need a clear audio track?

Yes - the transcript reflects the quality of the audio. A video recorded in a conference room with a distant microphone will produce more errors than one recorded with a headset or lapel mic. Picture quality is irrelevant; only the audio track is processed. If your recording was taken at a distance, moving a phone or small recorder closer to the speakers for the next session is the single most effective improvement.

Can I transcribe a video recorded in a language other than English?

Yes. The tool supports multiple languages. For short or noisy recordings, selecting the language explicitly rather than leaving it on auto-detect removes one source of error. For a bilingual recording, select the language that dominated the conversation.

What if my video has subtitles or captions already burned in?

Burned-in captions are part of the picture, not the audio. They do not affect transcription. The transcript is generated entirely from the audio track and will be independent of any on-screen text.

Can I transcribe a phone video recorded indoors?

Yes. Videos recorded on iPhone or Android are typically .mp4 or .mov files, both accepted directly. The main thing to check is file size: a long recording at high resolution can exceed 100 MB, in which case extracting the audio track first, as described above, brings it within limits while preserving every word of the conversation.

Transcribe a video file