How to convert MP3 to text free (with speaker labels)
2026-07-23
An MP3 from a 60-minute interview holds roughly 9,000 spoken words - about 30 pages of text. Typing that back manually takes 3 to 5 hours even at 60 words per minute. With WhoSayAh the same file returns a labelled transcript in a few minutes, free, with no account and nothing to install. This guide walks through every step and explains how to get the cleanest result from whatever MP3 you have.
Step-by-step: upload your MP3 and get a transcript
- Open the WhoSayAh home page and drop your MP3 onto it, or click to browse for the file. The upload starts immediately and a dialog appears showing the file name, size and detected length.
- Check that the detected length looks right. Files up to 100 MB and 60 minutes are accepted; if yours is longer, see the section below on splitting recordings.
- Set the language. Leave it on Detect automatically for most recordings. If the MP3 is very short (under 2 minutes) or recorded in a noisy environment, naming the language explicitly removes one source of error.
- Switch on Label speakers if more than one person is talking. Enter the number of speakers if you know it, or press I don't know - detect it. For a two-person interview, entering 2 produces consistently tighter attribution than leaving detection open.
- Press Transcribe. Lines appear as they are processed, so you can start reading before the file is fully done.
- Once the transcript finishes, press Name speakers. Each voice is shown with how long it spoke and its first line - enough to identify who is who from a meeting you attended. Replace Speaker 1, Speaker 2 with real names and the labels update across the entire transcript instantly.
- Copy to clipboard or download as
.txtor.vtt.
Free, no account, no watermark on the output - up to 100 MB and 60 minutes per file.
Which MP3 quality settings give the best transcript
Recording quality is the single biggest factor in transcript accuracy, more than any setting you choose in the tool. The decisions that matter most happen before you press record.
- Bitrate: 128 kbps or higher. Below 128 kbps, lossy compression starts to blur consonants - exactly the sounds that distinguish similar-sounding words. At 128 kbps and above, MP3 is entirely adequate for transcription. 192 kbps and 256 kbps are fine but rarely make a measurable difference to the output.
- One speaker per channel where possible. If your recording software can capture each participant on a separate track and export them as a stereo MP3 (interviewer left, guest right), speaker attribution becomes nearly perfect because the voices never truly overlap on the same channel.
- Keep background noise low. A café recording at 256 kbps will transcribe worse than a quiet room recorded at 96 kbps. Physical environment matters more than bitrate once you are above the 128 kbps floor.
- Do not re-encode. If you already have the recording in a lossless format (wav or flac), upload that instead. Each lossy conversion round-trips through compression and loses a little more detail. Start from the highest-quality source you have and let WhoSayAh handle the rest.
- Avoid noise reduction pre-processing. Aggressive noise-reduction plugins can introduce artefacts that are harder to parse than the original background noise. Upload the raw export unless the noise is so severe the recording is barely intelligible.
A general rule: if the recording is clear to your own ear, the transcript will be accurate. Where you can catch words through conscious effort but the audio is muddy, expect some errors on short interjections and words that share a consonant cluster.
MP3 vs other audio formats: what WhoSayAh accepts
MP3 is the most common format for exported podcast recordings, Zoom local files and voice memo apps, so it is the one most people upload. But WhoSayAh accepts a wider range, which can save you a conversion step if your recorder exports something else.
| Format | Accepted | Notes |
|---|---|---|
| mp3 | Yes | Most common; 128 kbps or above recommended |
| m4a | Yes | Default from iPhone Voice Memos and many recorders |
| wav | Yes | Lossless; larger files, no quality penalty |
| flac | Yes | Lossless compression; smaller than wav at identical quality |
| ogg / opus | Yes | Common from browser recordings and OBS |
| aac | Yes | Used by some video-editing exports |
| mp4 / mov / mkv / webm | Yes | Video files; only the audio track is read |
Video containers are accepted as-is - you do not need to strip the audio before uploading. If you recorded a Google Meet or Zoom call and the export is an mp4, drop the mp4 directly. The size limit is 100 MB and the length limit is 60 minutes per file. A 60-minute MP3 at 128 kbps weighs around 55 MB, which is well within the limit.
Longer recordings: how to split a file over the limit
A full-day conference recording or a long deposition may exceed the per-file limits. The practical fix is to split the file into segments under 60 minutes each before uploading. Free tools that handle this without re-encoding include:
- Audacity - open the file, set the time selection to the segment boundary, and export. Audacity can export a range directly to MP3 without touching the rest of the file. Available on Windows, macOS and Linux.
- mp3splt - a command-line tool that cuts MP3 files at a time marker without decoding and re-encoding, so there is no quality loss at all. One command per segment; runs on any platform.
- VLC - Record menu lets you capture a section of a file to a new file, though the UI is slightly awkward for precise boundaries.
When splitting a multi-person recording, cut at a silence or a natural pause rather than mid-sentence. Speaker attribution across a segment boundary restarts at Speaker 1 - name the speakers in each segment separately and reconcile when combining the transcripts.
Speaker labels: what they are and why they matter
A plain-text transcript of a two-person conversation is nearly unusable for reference: you can read every word and still not know who made the commitment, who raised the objection, or who gave the number. Speaker labels attach each line to the voice that spoke it, automatically, turning a wall of text into a structured record.
The labels arrive as Speaker 1, Speaker 2 and so on.
After transcription, Name speakers shows each speaker alongside how long
they spoke and their first line - for an interview you conducted, that is usually enough
to identify who is who at a glance. Enter real names and they propagate across every line
of the transcript instantly, including the downloaded .txt and
.vtt files.
Attribution accuracy is highest when the speaker count is entered exactly. Leaving the tool to estimate works well for most recordings, but a wide-pitched voice or someone who changes pace sharply between calm and animated turns can register as 2 speakers instead of 1. If you know how many people are in the recording - and for a structured interview you almost always do - enter the count. For recordings where the number genuinely varies, I don't know - detect it handles it reasonably. For a deeper look at two-person recordings specifically, the interview transcription guide covers volume imbalance, bilingual conversations and other edge cases.
Download formats: .txt and .vtt
Once the transcript is ready you have three ways to take it:
- Copy to clipboard - one click, the full transcript (speaker names and all) lands in your clipboard for pasting directly into a document, CRM, or email.
- .txt - plain text, one line per speaking turn, speaker name prefixed. Opens in any text editor. Paste into meeting notes, a research document, or a ticket. No proprietary format, no dependency on any software staying available.
- .vtt - WebVTT with start and end timestamps on every cue. Drop this
directly into an HTML5
<video>element, upload it to YouTube as a caption file, or import it into Premiere Pro, DaVinci Resolve or any editor that accepts external captions. No conversion step required; WebVTT is a W3C standard.
For teams that publish a large volume of recorded content - weekly call recordings, podcast episodes, training videos - manually uploading one file at a time is not practical at scale. CrocOTT handles the VOD platform side of that workflow, keeping caption files synced with a growing catalogue and removing the per-file manual step.
What happens to your MP3 after upload
The uploaded file is deleted as soon as the transcript is ready. The transcript itself is deleted after 24 hours. There is no account, so nothing links either to you personally. The privacy policy states exactly what is stored and for how long; the terms of use set out the per-file size and length limits in full.
If the recording contains material that cannot leave your network - legal depositions, medical consultations, NDAs, anything subject to a data-residency requirement - the alternative is a self-hosted installation where the audio never reaches an external server. FastoCloud provides that infrastructure: processing runs inside your own environment, there is no per-minute fee, and the workflow for the person uploading the file is identical to what is described above.
Frequently asked questions
Does my MP3 have to be a specific bitrate?
No minimum is enforced at upload. In practice, 128 kbps or above gives the cleanest results. Files recorded at lower bitrates - some voice-memo apps default to 64 kbps or 32 kbps - will still transcribe, but you may notice more errors on consonant clusters and very short words where the compression has blurred the signal.
How long does transcription take?
Typically a few minutes for a 60-minute file. The tool is asynchronous: you upload once, get a job ID, and lines appear as they are processed. You do not need to keep the tab focused - come back when it is done.
Can I convert a video file instead of an MP3?
Yes. mp4, mov, mkv and webm are all accepted. Only the audio track is read. Upload the video file directly; you do not need to extract the audio first.
What if I do not know how many speakers are in the recording?
Press I don't know - detect it. The tool estimates the speaker count from the audio. Results are good for most recordings; where a speaker has a very wide pitch range, the estimate may be one too high.
Is there a limit per day or per account?
There is no account. The limits are per upload: 100 MB and 60 minutes per file. You can upload as many files as you like, one at a time.