How to convert MP3 to text free (with speaker labels)

An MP3 from a 60-minute interview holds roughly 9,000 spoken words - about 30 pages of text. Typing that back manually takes 3 to 5 hours even at 60 words per minute. With WhoSayAh the same file returns a labelled transcript in a few minutes, free, with no account and nothing to install. This guide walks through every step and explains how to get the cleanest result from whatever MP3 you have.

Step-by-step: upload your MP3 and get a transcript

  1. Open the WhoSayAh home page and drop your MP3 onto it, or click to browse for the file. The upload starts immediately and a dialog appears showing the file name, size and detected length.
  2. Check that the detected length looks right. Files up to 100 MB and 60 minutes are accepted; if yours is longer, see the section below on splitting recordings.
  3. Set the language. Leave it on Detect automatically for most recordings. If the MP3 is very short (under 2 minutes) or recorded in a noisy environment, naming the language explicitly removes one source of error.
  4. Switch on Label speakers if more than one person is talking. Enter the number of speakers if you know it, or press I don't know - detect it. For a two-person interview, entering 2 produces consistently tighter attribution than leaving detection open.
  5. Press Transcribe. Lines appear as they are processed, so you can start reading before the file is fully done.
  6. Once the transcript finishes, press Name speakers. Each voice is shown with how long it spoke and its first line - enough to identify who is who from a meeting you attended. Replace Speaker 1, Speaker 2 with real names and the labels update across the entire transcript instantly.
  7. Copy to clipboard or download as .txt or .vtt.

Free, no account, no watermark on the output - up to 100 MB and 60 minutes per file.

Which MP3 quality settings give the best transcript

Recording quality is the single biggest factor in transcript accuracy, more than any setting you choose in the tool. The decisions that matter most happen before you press record.

A general rule: if the recording is clear to your own ear, the transcript will be accurate. Where you can catch words through conscious effort but the audio is muddy, expect some errors on short interjections and words that share a consonant cluster.

MP3 vs other audio formats: what WhoSayAh accepts

MP3 is the most common format for exported podcast recordings, Zoom local files and voice memo apps, so it is the one most people upload. But WhoSayAh accepts a wider range, which can save you a conversion step if your recorder exports something else.

FormatAcceptedNotes
mp3YesMost common; 128 kbps or above recommended
m4aYesDefault from iPhone Voice Memos and many recorders
wavYesLossless; larger files, no quality penalty
flacYesLossless compression; smaller than wav at identical quality
ogg / opusYesCommon from browser recordings and OBS
aacYesUsed by some video-editing exports
mp4 / mov / mkv / webmYesVideo files; only the audio track is read

Video containers are accepted as-is - you do not need to strip the audio before uploading. If you recorded a Google Meet or Zoom call and the export is an mp4, drop the mp4 directly. The size limit is 100 MB and the length limit is 60 minutes per file. A 60-minute MP3 at 128 kbps weighs around 55 MB, which is well within the limit.

Longer recordings: how to split a file over the limit

A full-day conference recording or a long deposition may exceed the per-file limits. The practical fix is to split the file into segments under 60 minutes each before uploading. Free tools that handle this without re-encoding include:

When splitting a multi-person recording, cut at a silence or a natural pause rather than mid-sentence. Speaker attribution across a segment boundary restarts at Speaker 1 - name the speakers in each segment separately and reconcile when combining the transcripts.

Speaker labels: what they are and why they matter

A plain-text transcript of a two-person conversation is nearly unusable for reference: you can read every word and still not know who made the commitment, who raised the objection, or who gave the number. Speaker labels attach each line to the voice that spoke it, automatically, turning a wall of text into a structured record.

The labels arrive as Speaker 1, Speaker 2 and so on. After transcription, Name speakers shows each speaker alongside how long they spoke and their first line - for an interview you conducted, that is usually enough to identify who is who at a glance. Enter real names and they propagate across every line of the transcript instantly, including the downloaded .txt and .vtt files.

Attribution accuracy is highest when the speaker count is entered exactly. Leaving the tool to estimate works well for most recordings, but a wide-pitched voice or someone who changes pace sharply between calm and animated turns can register as 2 speakers instead of 1. If you know how many people are in the recording - and for a structured interview you almost always do - enter the count. For recordings where the number genuinely varies, I don't know - detect it handles it reasonably. For a deeper look at two-person recordings specifically, the interview transcription guide covers volume imbalance, bilingual conversations and other edge cases.

Download formats: .txt and .vtt

Once the transcript is ready you have three ways to take it:

For teams that publish a large volume of recorded content - weekly call recordings, podcast episodes, training videos - manually uploading one file at a time is not practical at scale. CrocOTT handles the VOD platform side of that workflow, keeping caption files synced with a growing catalogue and removing the per-file manual step.

What happens to your MP3 after upload

The uploaded file is deleted as soon as the transcript is ready. The transcript itself is deleted after 24 hours. There is no account, so nothing links either to you personally. The privacy policy states exactly what is stored and for how long; the terms of use set out the per-file size and length limits in full.

If the recording contains material that cannot leave your network - legal depositions, medical consultations, NDAs, anything subject to a data-residency requirement - the alternative is a self-hosted installation where the audio never reaches an external server. FastoCloud provides that infrastructure: processing runs inside your own environment, there is no per-minute fee, and the workflow for the person uploading the file is identical to what is described above.

Frequently asked questions

Does my MP3 have to be a specific bitrate?

No minimum is enforced at upload. In practice, 128 kbps or above gives the cleanest results. Files recorded at lower bitrates - some voice-memo apps default to 64 kbps or 32 kbps - will still transcribe, but you may notice more errors on consonant clusters and very short words where the compression has blurred the signal.

How long does transcription take?

Typically a few minutes for a 60-minute file. The tool is asynchronous: you upload once, get a job ID, and lines appear as they are processed. You do not need to keep the tab focused - come back when it is done.

Can I convert a video file instead of an MP3?

Yes. mp4, mov, mkv and webm are all accepted. Only the audio track is read. Upload the video file directly; you do not need to extract the audio first.

What if I do not know how many speakers are in the recording?

Press I don't know - detect it. The tool estimates the speaker count from the audio. Results are good for most recordings; where a speaker has a very wide pitch range, the estimate may be one too high.

Is there a limit per day or per account?

There is no account. The limits are per upload: 100 MB and 60 minutes per file. You can upload as many files as you like, one at a time.

Convert your MP3 now