How to transcribe audio to text for free
2026-07-19
A one-hour meeting recording takes around 4 hours to transcribe by hand at average typing speed. A 20-minute interview takes roughly 80 minutes. Neither task is necessary any more: upload the file to WhoSayAh, wait a few minutes, and copy the text. Free, no account, nothing to install. This guide explains how to do that and how to get the best result from whatever recording you have.
The quickest route
- Open the home page and drop your file on it. It starts uploading immediately and a window shows the file's name, size and detected length.
- Leave the language on Detect automatically unless the recording is very short or noisy, in which case naming the language helps.
- Keep Label speakers on if more than one person is talking, and enter how many. If you are not sure, press I don't know - detect it.
- Press Transcribe and read lines as they appear.
- Once the transcript is complete, press Name speakers to replace
Speaker 1, Speaker 2 with real names, then copy the transcript or
download as
.txtor.vtt.
Free, up to 100 MB and 120 minutes per file, no signup.
What you can upload
Most common audio and video formats are accepted. Video containers do not need to be converted first - only the audio track is read, so a screen recording or video call file uploads as-is.
| Category | Accepted formats |
|---|---|
| Audio | mp3, m4a, wav, ogg, opus, flac, aac |
| Video | mp4, mov, mkv, webm |
The limits are 100 MB and 120 minutes per file. Both are checked at upload time, before any transcription work begins, so if a file is too large you find out within seconds rather than after a long wait. If you have a longer recording - a full deposition, a conference day - split it into segments under the limit. Most free audio editors, including Audacity, can cut a file by time marker without re-encoding.
How to get a better transcript
Transcription quality is largely decided before you upload. In rough order of impact:
- Record close to the speaker. A microphone 2 feet away produces far cleaner audio than one placed across a room. For remote calls, a headset or external USB microphone is the single most effective improvement: it captures your own voice clearly and avoids room echo and keyboard noise.
- Avoid overlapping speech. Crosstalk is the hardest problem for both word accuracy and speaker attribution. Even 2 seconds of two voices at once can shift the next line to the wrong speaker. If you are running an interview or a panel, gentle turn-taking cues - a brief pause, "go ahead" - help significantly.
- Name the language if the recording is short or noisy. Auto-detection works from the audio signal and has less to go on when the recording is under a minute or the environment is loud. Naming the language explicitly removes one source of error in both cases.
- Do not pre-compress. A file that has already been through multiple lossy conversions has lost detail the transcriber could have used. If you have a choice between a raw recording and a heavily compressed copy, use the raw version.
- Prefer higher-quality formats. Among lossless options, wav and flac carry the most information. Among lossy ones, mp3 at 128 kbps or above, m4a, and opus are all fine. Format matters far less than recording conditions, but when given the choice prefer the higher-quality file.
A recording that is clear to your own ear will produce a clean transcript. One that is hard to follow will still produce one, but expect more errors on very short interjections and anywhere voices overlap.
Why speaker labels matter
For a single voice - a lecture, a podcast solo, a voice memo - plain text is fine. For anything with two or more people, a transcript without speaker labels is close to unusable: you can read every word and still not know who committed to what, who asked the question, or who raised the objection.
Speaker labels turn a transcript into a record. Each line is attributed to the person who said it, automatically, without you having to listen and label by hand. For a 1-hour meeting, that saves roughly 3 to 4 hours of manual work compared with listening back and annotating yourself.
The labels arrive as Speaker 1, Speaker 2 and so on.
Once the transcript is complete, Name speakers lets you replace those with
real names. Each voice is listed with how long it spoke and its first line - for a
meeting you have just attended, that is usually enough to identify who is who. The names
apply across the entire transcript instantly and go into the downloaded .txt
and .vtt files.
Speaker labelling works best when the speaker count is entered exactly. Leaving the tool to estimate can cause a wide-pitched voice, or someone who changes pace noticeably between calm and animated turns, to be counted as 2 people. If you know the count, enter it; if not, I don't know - detect it gives a reasonable result.
For a deeper look at the specific settings that matter most for a two-person recording, the interview transcription guide walks through each step in detail, including how to handle volume imbalance and bilingual conversations.
Output formats: .txt and .vtt
Once the transcript is complete, you have three options:
- Copy to clipboard - one click, the full transcript lands in your clipboard ready to paste into a document, email, or ticket.
- .txt download - plain text, one line per speaking turn, speaker name prefixed. Use this to paste into meeting notes, share with a colleague, or attach to a project record. No proprietary format, opens in any text editor.
- .vtt download - WebVTT with precise start and end timestamps on
each cue. Drop this directly into an HTML5
<video>element as a subtitle track, upload it to YouTube as a caption file, or import it into video-editing software that accepts external captions. No conversion step needed.
WebVTT is a web standard, so the .vtt file works across tools without any
additional processing. If you are publishing a recorded meeting or interview as video
content, the file is ready to use immediately. Teams working with a large library of
recorded content at scale - dozens of call recordings a week - typically
automate this workflow. CrocOTT
handles the VOD platform side of that, keeping captions synced with the catalogue and
removing the need to upload files one by one.
What happens to your file
The uploaded file is deleted as soon as the transcript is ready. The transcript is deleted after 24 hours. There is no account, so nothing links either to you. The privacy policy spells out exactly what is stored and for how long, and the terms of use list the size and length limits in full.
For recordings that cannot leave your own network - client calls under NDA, medical or legal material, anything subject to a data-residency regulation - the alternative is a self-hosted installation where the audio never reaches an external server. FastoCloud provides the infrastructure for that: the processing runs inside your own environment, there is no per-minute charge, and the workflow for whoever uploads the file is identical to what is described above.
Frequently asked questions
Can I transcribe audio without installing software?
Yes. Upload the file in your browser and the transcript comes back on the same page. Nothing to install, no account required. The tool runs in any modern browser on any operating system.
What file formats does WhoSayAh accept?
Audio: mp3, m4a, wav, ogg, opus, flac, aac. Video: mp4, mov, mkv, webm. Video files do not need to be stripped to audio first - only the audio track is read. The size limit is 100 MB and the length limit is 120 minutes per file.
Can it tell speakers apart?
Yes. With speaker labels enabled, each line is attributed to a speaker. The tool works automatically - you do not need separate microphone tracks. Entering the exact speaker count gives the best results; if you do not know the count, the auto-detect option provides a reasonable estimate.
How accurate is the transcript?
Accuracy depends mainly on recording quality. A clear recording with one speaker per turn, recorded close to a microphone, produces a very accurate transcript in most languages. Crosstalk and background noise are the main causes of errors; very short interjections (a single-word "yeah") are the lines most likely to be attributed to the wrong speaker.
Is transcription really free?
Yes. There is no charge, no account, and no watermark on the output. The only constraints are the file size and length limits, which apply per upload.