How to transcribe a video file to text for free (with speaker labels)
2026-08-14
An hour-long meeting recorded as an mp4 takes roughly 4 hours to transcribe by hand - and that is before you try to work out who said what. Most tools that accept video either strip the audio and charge per minute, or lock transcription behind a subscription. WhoSayAh transcribes video files directly in your browser: no conversion, no account, no charge for files within the size and length limits, and automatic speaker labels so every line is attributed to the person who said it.
What video formats are accepted
Common video containers upload as-is. Only the audio track is read during transcription, so a 500 MB screen recording of a 30-minute call produces the same result as an audio-only export of the same call - without any extra conversion step.
| Category | Accepted formats | Typical size for 30 minutes |
|---|---|---|
| Video | mp4, mov, mkv, webm | 200-600 MB |
| Audio | mp3, m4a, wav, ogg, opus, flac, aac | 20-60 MB |
The upload limit is 100 MB and 120 minutes per file. Both are checked immediately at upload time, so an over-sized file is rejected in seconds rather than after a long wait. If you have a longer recording - a full conference session, a deposition, a day of interviews - split it into segments under the limit. Most free video editors can cut by time marker without re-encoding.
Video files frequently exceed the size limit even for moderate durations, because a screen
recording at high resolution can run to several hundred megabytes for a single hour. The
quickest fix is to extract the audio track before uploading. On a Mac, QuickTime Player exports
audio only via File → Export → Audio Only. On Windows and Linux, the free VLC media player
converts a video to .m4a in under a minute via Media → Convert/Save, selecting an
audio-only profile. The transcript is identical; the file is a fraction of the size.
Step by step: from video file to labelled transcript
- Open the home page and drop your file. A dialog appears showing the file name, size and detected length. If the file is over 100 MB or 120 minutes, the dialog says so immediately.
- Confirm the language. Leave Detect automatically on for most recordings. Select the language explicitly when the recording is under 2 minutes or the environment was noisy - auto-detection has less to go on in those cases.
- Set the speaker count. Keep Label speakers on and enter how many people actually spoke - not how many attended. A 12-person all-hands where 4 people contribute is a 4-speaker recording. If you are not sure, press I don't know - detect it.
- Press Transcribe and read along. Lines appear on screen as they are processed. A 20-minute recording is typically readable within 2 to 3 minutes; a full hour within 6 to 8 minutes. You do not need to stay on the page, but results appear there.
- Name the speakers. When the transcript is complete, press Name speakers. Each voice is listed with the total time it spoke and its first line. For a meeting you just attended, placing everyone takes about 30 seconds. The names replace Speaker 1, Speaker 2 across the entire transcript in one step.
-
Download or copy. Take the transcript as
.txt(plain text, one line per speaking turn, name prefixed) or.vtt(WebVTT with timestamps, ready for video software or YouTube). Or copy the whole thing to clipboard in one click.
Why you do not need to convert the video first
Many transcription tools accept video in name only: they re-encode the file server-side before processing it, adding a format-conversion step you do not see. WhoSayAh reads the audio track directly from the container, skipping that conversion. In practice:
- The file goes up once. There is no second queue for a format conversion between upload and transcription.
- No quality loss from re-encoding. Running lossy audio through a second encode introduces a second generation of compression artefacts. Reading the original track avoids that.
- Fewer steps. Drag the file from your downloads folder and you are done. Whether it is mp4, mov, mkv or webm does not matter.
The one case where pre-conversion still makes sense is when the video file is very large. A 2-hour screen recording at 1080p is often 1.5-3 GB, which will exceed 100 MB. There, the right move is to extract the audio - not to compress the video. The audio track from that same recording is typically 50-120 MB.
Speaker labels: how they work and when they help most
A transcript without speaker labels is a wall of text. You can read every word and still not know who made the commitment, who raised the objection, or who owns the follow-up. Speaker labels assign each line to the person who said it, automatically, without you listening back and annotating by hand. For a 1-hour recording, that saves 3 to 4 hours of work.
Speaker labels work from the audio content of the file, not from metadata or separate microphone tracks. This means they work with any video recording - a Zoom mp4, a Teams meeting export, a phone video of an in-person session. Attribution is strongest when:
- Speakers do not talk at the same time. Two voices overlapping for even 2 seconds can push the following line to the wrong speaker. Turn-taking discipline that good meetings need anyway is exactly what transcription rewards.
- The speaker count is entered exactly. A wide-pitched voice, or someone who shifts between animated and quiet turns, can be counted as two people if the tool is left to estimate. If you know the count, enter it.
- Recording levels are roughly balanced. A recording where one participant is twice as loud as the others is harder to separate cleanly. Headsets and external microphones are the single most effective fix before the next session.
Once the transcript is complete, Name speakers lists each voice with how long
it spoke and its first line - enough to identify everyone in a meeting you just attended. The
names go into both the .txt and .vtt downloads. For a two-person
interview with a remote guest, the
interview transcription guide
covers the specific settings that matter most for that format, including how to handle volume
imbalance between an interviewer and a remote guest.
Output formats: .txt, .vtt and copy
Once the transcript is ready there are three ways to take the text:
- Copy to clipboard - one click, the full transcript lands ready to paste into a document, email, or ticket. Speaker names are included.
- .txt download - one line per speaking turn, name prefixed. Attach it to a project record, share it with a colleague, or feed it into a summariser or search index.
- .vtt download - WebVTT with precise timestamps on each cue. Drop this file
into an HTML5
<video>element, upload it to YouTube as captions, or import it into editing software that accepts external subtitle tracks. The subtitles guide walks through each destination in detail, from YouTube to your own website.
WebVTT is a web standard, so the .vtt file works across platforms without any
additional tooling or conversion. Teams that publish recorded video at scale - a legal firm
archiving depositions, a media company captioning interviews, an HR department releasing
training sessions - reach a point where uploading one file at a time is not practical.
CrocOTT handles the
VOD-platform side of that: captions synced across a full catalogue, accessibility compliance
built in, no manual upload loop.
What happens to your video file
The uploaded file is deleted as soon as the transcript is ready. The transcript is deleted after 24 hours. There is no account, so nothing links either to you. The privacy policy states exactly what is stored and for how long. The terms of use carry the upload size and length limits in full.
For recordings that cannot leave your own network - NDA-covered client calls, medical consultations, legal depositions, anything subject to a data-residency regulation - the alternative is a self-hosted installation where the audio never reaches an external server. FastoCloud provides the infrastructure for that: the same browser-based upload experience described in this guide, running entirely inside your own environment with no per-minute charge.
Common questions
Does the video need a clear audio track?
Yes - the transcript reflects the quality of the audio. A video recorded in a conference room with a distant microphone will produce more errors than one recorded with a headset or lapel mic. Picture quality is irrelevant; only the audio track is processed. If your recording was taken at a distance, moving a phone or small recorder closer to the speakers for the next session is the single most effective improvement.
Can I transcribe a video recorded in a language other than English?
Yes. The tool supports multiple languages. For short or noisy recordings, selecting the language explicitly rather than leaving it on auto-detect removes one source of error. For a bilingual recording, select the language that dominated the conversation.
What if my video has subtitles or captions already burned in?
Burned-in captions are part of the picture, not the audio. They do not affect transcription. The transcript is generated entirely from the audio track and will be independent of any on-screen text.
Can I transcribe a phone video recorded indoors?
Yes. Videos recorded on iPhone or Android are typically .mp4 or .mov
files, both accepted directly. The main thing to check is file size: a long recording at high
resolution can exceed 100 MB, in which case extracting the audio track first, as
described above, brings it within limits while preserving every word of the conversation.