How to transcribe an interview with speaker labels

An interview recording with two voices and no speaker labels is one of the most frustrating documents to work from. You can hear every word and still not know who asked the question, who gave the answer, or who made the commitment. Turning it into something readable without speaker attribution means listening and typing at the same time - work that takes at least as long as the interview itself.

WhoSayAh does both steps in the browser: audio to text, then each line attributed to a named speaker. No account, no software to install, free for files up to 100 MB and 60 minutes. The steps below cover the whole process from raw file to a named transcript you can copy or download as .txt or .vtt.

Before you upload: what makes a good interview recording

Speaker labelling quality is decided before you open the tool. Three things have the most impact:

A recording that is clear to your own ear will produce a clean transcript. One that is difficult to follow will still produce one, but expect more errors on very short interjections and anywhere voices overlap.

Step by step

1. Upload the file

Open the WhoSayAh home page and drop the interview file onto it, or click to browse. The file starts uploading immediately. A window opens showing the file's name, size and detected length.

If the file exceeds 100 MB or 60 minutes, a rejection message appears at this step - not after you have waited for transcription. That early check means you can trim or split the file before spending any time waiting.

2. Set the language

Leave the language on Detect automatically unless the recording is very short (under a minute) or particularly noisy. In both cases there is less audio to work from, and naming the language explicitly helps. For bilingual interviews where speakers switch languages mid-conversation, choose the language used for the larger portion.

3. Enable speaker labels and enter the count

This is the most important step for an interview. Switch Label speakers on. In the How many people are speaking? field, enter the exact count. For a one-to-one interview the answer is 2; for a panel it is the number of people who actually spoke.

Entering the correct count is the single setting that most improves results. Left to estimate, the tool can over-split a voice - treating one person as two if they have a wide pitch range or change pace noticeably between calm answers and animated ones. If you genuinely do not know the count, I don't know - detect it will estimate from the audio.

4. Press Transcribe and read as it arrives

Press Transcribe. Lines appear as they are processed, so you can start reading before the job finishes. A twenty-minute interview is typically readable within a couple of minutes of pressing start; longer files take proportionally longer.

Each line is prefixed with Speaker 1, Speaker 2 and so on. At this stage the tool knows the voices are distinct but has no way to know whose name belongs to which.

5. Name the speakers

Once the transcript is complete, press Name speakers. A panel lists each voice with two pieces of information: how long it spoke, and its first line. For an interview that is almost always enough - the interviewer tends to ask the opening question and has the shorter share of the runtime; the subject has the longer. Type in the real names and they are applied across the entire transcript instantly.

6. Copy or download

Copy the full text with one click, or download in two formats:

What the output looks like

A named interview transcript in .txt format reads like this:

TimeSpeakerLine
0:00 Interviewer So tell me how you got started with this.
0:05 Guest It actually begins with a problem I kept running into in 2019...
0:22 Interviewer And what changed after that?
0:26 Guest Everything, really. The constraints were completely different...

The .vtt version adds precise timestamps on every cue rather than a human-readable offset. That makes it suitable for an HTML5 <video> subtitle track, for YouTube caption upload, and for podcast platforms that support transcript files alongside the audio.

If you are publishing the interview as a video or podcast episode, the .vtt file is ready to use without any conversion step. Operators captioning a larger library of recorded interviews and conference talks at scale - dozens or hundreds of files - typically automate this workflow. CrocOTT handles the VOD platform side of that, removing the need to upload each file by hand and keeping the captions synced with the catalogue.

Why interviews are harder than single-voice recordings

A lecture or voice memo has one speaker, so there is nothing to separate. An interview introduces challenges that single-voice recordings do not have.

Short interjections. A one-word "right" or "yes" carries very little voice information. Those are the lines most likely to be attributed to the wrong speaker. Complete sentences are reliable; very short turns are not. If attribution accuracy matters for every single line, the cleanest fix is recording the interview with each person on a separate track - most podcast recording tools offer this as dual mono or multi-track mode. Either track can be uploaded separately, and the transcripts merged afterwards.

Volume imbalance. Phone interviews commonly suffer from this: the remote caller comes through at a lower level than the local speaker. The quieter voice is more likely to produce transcription errors and mis-attribution. Some call-recording apps compensate automatically; if yours does not, normalising the audio in a free editor before uploading helps. Matching levels between speakers to within about 3 dB is sufficient.

Knowing the speaker count matters more than anything else. The two-minute effort of entering the correct number in step 3 above consistently produces better output than any post-processing step. If you do not know the count, start with a guess of 2 for a one-to-one interview and check whether the result looks right.

For more general advice on getting a clean transcript from any recording, the audio transcription guide covers microphone technique and format choice in more detail.

Privacy and your interview recordings

Recording an interview means recording another person. In many jurisdictions that requires explicit consent before you press record; everywhere else it is simply the right thing to do. Tell your subject, get their agreement, and document it if the recording will be published or used in a professional context.

On WhoSayAh's side: the uploaded file is deleted as soon as the transcript is ready. The transcript is deleted after 24 hours. There is no account, so nothing links either to you. The privacy policy states exactly what is kept and for how long. The terms of use list the file size and length limits in full.

For interviews that cannot leave your own network - client conversations under NDA, medical or legal material, anything subject to a data-residency rule - the alternative is a self-hosted installation where the audio never reaches an external server. FastoCloud provides the infrastructure for that use case: the processing stays inside your own environment, and there is no per-minute charge. The workflow for the end user is identical to what is described above.

Transcribe an interview