How to transcribe audio to text for free

A one-hour meeting recording takes around 4 hours to transcribe by hand at average typing speed. A 20-minute interview takes roughly 80 minutes. Neither task is necessary any more: upload the file to WhoSayAh, wait a few minutes, and copy the text. Free, no account, nothing to install. This guide explains how to do that and how to get the best result from whatever recording you have.

The quickest route

  1. Open the home page and drop your file on it. It starts uploading immediately and a window shows the file's name, size and detected length.
  2. Leave the language on Detect automatically unless the recording is very short or noisy, in which case naming the language helps.
  3. Keep Label speakers on if more than one person is talking, and enter how many. If you are not sure, press I don't know - detect it.
  4. Press Transcribe and read lines as they appear.
  5. Once the transcript is complete, press Name speakers to replace Speaker 1, Speaker 2 with real names, then copy the transcript or download as .txt or .vtt.

Free, up to 100 MB and 120 minutes per file, no signup.

What you can upload

Most common audio and video formats are accepted. Video containers do not need to be converted first - only the audio track is read, so a screen recording or video call file uploads as-is.

CategoryAccepted formats
Audiomp3, m4a, wav, ogg, opus, flac, aac
Videomp4, mov, mkv, webm

The limits are 100 MB and 120 minutes per file. Both are checked at upload time, before any transcription work begins, so if a file is too large you find out within seconds rather than after a long wait. If you have a longer recording - a full deposition, a conference day - split it into segments under the limit. Most free audio editors, including Audacity, can cut a file by time marker without re-encoding.

How to get a better transcript

Transcription quality is largely decided before you upload. In rough order of impact:

A recording that is clear to your own ear will produce a clean transcript. One that is hard to follow will still produce one, but expect more errors on very short interjections and anywhere voices overlap.

Why speaker labels matter

For a single voice - a lecture, a podcast solo, a voice memo - plain text is fine. For anything with two or more people, a transcript without speaker labels is close to unusable: you can read every word and still not know who committed to what, who asked the question, or who raised the objection.

Speaker labels turn a transcript into a record. Each line is attributed to the person who said it, automatically, without you having to listen and label by hand. For a 1-hour meeting, that saves roughly 3 to 4 hours of manual work compared with listening back and annotating yourself.

The labels arrive as Speaker 1, Speaker 2 and so on. Once the transcript is complete, Name speakers lets you replace those with real names. Each voice is listed with how long it spoke and its first line - for a meeting you have just attended, that is usually enough to identify who is who. The names apply across the entire transcript instantly and go into the downloaded .txt and .vtt files.

Speaker labelling works best when the speaker count is entered exactly. Leaving the tool to estimate can cause a wide-pitched voice, or someone who changes pace noticeably between calm and animated turns, to be counted as 2 people. If you know the count, enter it; if not, I don't know - detect it gives a reasonable result.

For a deeper look at the specific settings that matter most for a two-person recording, the interview transcription guide walks through each step in detail, including how to handle volume imbalance and bilingual conversations.

Output formats: .txt and .vtt

Once the transcript is complete, you have three options:

WebVTT is a web standard, so the .vtt file works across tools without any additional processing. If you are publishing a recorded meeting or interview as video content, the file is ready to use immediately. Teams working with a large library of recorded content at scale - dozens of call recordings a week - typically automate this workflow. CrocOTT handles the VOD platform side of that, keeping captions synced with the catalogue and removing the need to upload files one by one.

What happens to your file

The uploaded file is deleted as soon as the transcript is ready. The transcript is deleted after 24 hours. There is no account, so nothing links either to you. The privacy policy spells out exactly what is stored and for how long, and the terms of use list the size and length limits in full.

For recordings that cannot leave your own network - client calls under NDA, medical or legal material, anything subject to a data-residency regulation - the alternative is a self-hosted installation where the audio never reaches an external server. FastoCloud provides the infrastructure for that: the processing runs inside your own environment, there is no per-minute charge, and the workflow for whoever uploads the file is identical to what is described above.

Frequently asked questions

Can I transcribe audio without installing software?

Yes. Upload the file in your browser and the transcript comes back on the same page. Nothing to install, no account required. The tool runs in any modern browser on any operating system.

What file formats does WhoSayAh accept?

Audio: mp3, m4a, wav, ogg, opus, flac, aac. Video: mp4, mov, mkv, webm. Video files do not need to be stripped to audio first - only the audio track is read. The size limit is 100 MB and the length limit is 120 minutes per file.

Can it tell speakers apart?

Yes. With speaker labels enabled, each line is attributed to a speaker. The tool works automatically - you do not need separate microphone tracks. Entering the exact speaker count gives the best results; if you do not know the count, the auto-detect option provides a reasonable estimate.

How accurate is the transcript?

Accuracy depends mainly on recording quality. A clear recording with one speaker per turn, recorded close to a microphone, produces a very accurate transcript in most languages. Crosstalk and background noise are the main causes of errors; very short interjections (a single-word "yeah") are the lines most likely to be attributed to the wrong speaker.

Is transcription really free?

Yes. There is no charge, no account, and no watermark on the output. The only constraints are the file size and length limits, which apply per upload.

Try it with a file