Best Free Transcription Tools With Speaker Labels in 2026
2026-07-30
Speaker labels - the feature that attributes each line of a transcript to the person who said it - are available on nearly every transcription tool today, but "free with speaker labels" covers an enormous range. Otter.ai's free plan allows exactly 3 file imports for the lifetime of the account. Notta's free tier caps individual recordings at 3 minutes each. Several tools advertise speaker detection as free and then activate the paywall the moment you try to export. This comparison cuts through that: 6 tools measured against what you can actually do on the free tier, without a credit card, in 2026.
What "free with speaker labels" actually means
Three conditions must all be true for a transcription tool to be genuinely free with speaker labels: no credit card to start, speaker attribution included on the free tier, and a usable output format you can download without paying. A tool that meets all three is genuinely free. Most meet one or two.
The hidden cost that trips people most often is the file import limit. Meeting assistants are optimised for recurring paid subscriptions, so they throttle file uploads hard on free plans - imports are a loss-leader for bot seats, not the point of the product. Tools that take any recording you already have - a phone interview, a Zoom local recording, a webinar download - handle uploads more generously because the upload is the product. Understanding that split before you sign up saves the frustration of hitting a wall on your fourth recording.
A second axis worth checking is output format. A transcript you can only read on-screen, in a proprietary app, is significantly less useful than one you can download as plain text or as a subtitle file ready for a video editor. The table below includes both.
6 free transcription tools with speaker labels compared
| Tool | Speaker labels free | Free limit | File import limit | Export formats (free) | Account required |
|---|---|---|---|---|---|
| WhoSayAh | Yes | 100 MB / 60 min per file | Unlimited | .txt, .vtt | No |
| Otter.ai | Yes | 300 min/mo, 30 min per conversation | 3 imports - lifetime | .txt, .docx, .pdf | Yes |
| Fireflies.ai | Yes (limited credits) | Limited transcription credits | Very limited | On-screen only (free) | Yes |
| Notta | Yes | 120 min/mo; 3 min per recording | Included in monthly cap | Limited on free tier | Yes |
| Google Meet built-in | Yes (live calls only) | Requires paid Workspace plan | No file upload | Google Docs (auto-saved) | Workspace account |
| Self-hosted (FastoCloud) | Yes | Your own hardware, no per-minute charge | Unlimited | .txt, .vtt | Your own system |
1. WhoSayAh - free, no account, unlimited imports
WhoSayAh is built around a single job: take a recording you already have and return a labelled transcript. Drop an audio or video file (mp3, m4a, wav, ogg, opus, flac, aac, mp4, mov, mkv, webm) on the home page, tell it how many speakers were present - or let it detect automatically - and read the lines as they arrive. Files up to 100 MB and 60 minutes are free per upload, with no account and no import cap. There is no trial, no usage counter, and no monthly limit on the number of files.
Speaker lines come through in real time as the transcript builds, labelled Speaker 1,
Speaker 2 and so on. Once the job is complete, the Name speakers step
lets you replace those labels with real names - the change applies instantly across the entire
transcript and carries through into both export formats. Download as .txt for meeting
notes or documents, or as .vtt to attach directly to a video on YouTube, in an HTML5
player, or in a video editor that accepts external subtitle tracks, with no conversion needed.
Privacy terms are precise: the uploaded file is deleted as soon as the transcript is ready, and the transcript itself is deleted after 24 hours. There is no account, so nothing links either to you. The privacy policy states exactly what is retained and for how long - it is a legal commitment, not a marketing line. The terms of use list the file size and length limits in full. For a detailed walkthrough of the upload-to-download workflow, the step-by-step transcription guide covers every setting, including how to handle recordings that exceed the length limit.
2. Otter.ai - strong for live meetings, a bottleneck for file uploads
Otter.ai's free plan gives 300 minutes of transcription per month, capped at 30 minutes per conversation, with speaker attribution included. The catch that surprises most people is the file import limit: 3 imports in total, for the lifetime of the account - not per month. The first time you try to import a fourth recording, you reach the paywall.
Its Pro plan at $16.99 a month ($8.33 billed annually) raises the import limit to 10 per month - still a ceiling, at a monthly cost. Where Otter earns its keep is for teams that run recurring meetings: a bot joins Google Meet, Zoom or Teams calls, produces a searchable transcript with action items, and builds an archive over time. That workflow is where the 300 monthly minutes have room to breathe. For a folder of recordings you want to work through in one sitting, the 3-import ceiling is the main reason people arrive at alternatives lists.
3. Fireflies.ai - best option for meeting bots across platforms
Fireflies.ai is primarily a meeting assistant: it joins Zoom, Google Meet and Microsoft Teams calls, produces timestamped transcripts with summaries and action items, and makes past meetings searchable. Speaker labels are included. The free tier provides a real but limited pool of transcription credits that run out quickly on calls longer than 30 minutes. Paid plans start at $10 per user per month billed annually ($18 month-to-month); the Business tier at $19 per user per month per month removes the storage and credit caps that matter most for active teams.
Fireflies is the right pick if the job is recurring live calls across multiple platforms and you want a single archive of all of them. For files you already have, the credit constraints on the free tier make it a poor fit compared with tools designed around file upload.
4. Notta - worth considering for multilingual meetings
Notta's free tier covers 120 minutes of transcription per month in total, but individual recordings are capped at 3 minutes each - long enough to verify accuracy on a short clip and not much else. Speaker attribution is included across all plans. Where Notta has a genuine edge over the other tools on this list is multilingual support: Japanese, Chinese and Korean are handled notably well alongside English, which makes it the most practical free option when those language pairs appear in your calls. Paid plans begin at roughly $9 per month billed annually and lift the per-recording cap to a more usable 5 hours.
5. Google Meet built-in transcription
Google Meet can generate a transcript with speaker names for meetings you host - no bot participant in the call, no third-party account, no extra vendor. The transcript saves automatically to Google Drive after the call ends, complete with timestamps and attribution matched to each participant's Google profile. The limitation is access: the feature requires a paid Google Workspace plan (Business Standard or above, at $14 per user per month as of 2026), and it only covers live calls you host on Meet - not uploaded recordings from other platforms.
If your team already runs on a paid Workspace plan and your meetings are internal Google Meet calls, the built-in transcript is the cleanest option: no extra login, no file upload step, speaker names already resolved to real identities. For calls you attended but did not host, or recordings from Zoom or other platforms, you need a file tool alongside it. The Google Meet recording guide shows how to capture and transcribe those sessions.
6. Self-hosted transcription via FastoCloud
Every tool above, WhoSayAh included, processes your audio on someone else's server. For client calls under NDA, medical or legal recordings, or anything subject to a data-residency regulation, that can be disqualifying regardless of the privacy policy. The alternative is running transcription inside your own network, on your own hardware, where the audio never crosses your perimeter and there is no per-minute meter. FastoCloud provides the infrastructure for that: a Docker-based deployment behind your own nginx, with speaker labels and the same file-upload workflow for the people using it. There is no per-minute charge - only the cost of the hardware you already control.
Which tool fits which job
The right choice follows from where your recordings come from and what you need to do with them:
- Files you already have → WhoSayAh: no account, no import cap,
speaker labels, exports as
.txtor.vtt, free up to 100 MB and 60 minutes per file. - Live meetings you want a bot to join → Fireflies.ai for cross-platform (Zoom, Teams, Meet); Google Meet built-in if your organisation already pays for Workspace.
- Multilingual calls (English + Japanese, Chinese, Korean) → Notta, accepting the 3-minute per-recording cap on the free tier.
- Audio that cannot leave your network → Self-host on FastoCloud, where processing runs inside your own environment.
Teams working through a large library of recorded content - dozens of calls a week, or a growing
archive of interviews - often pair a free upload tool for first-pass transcription with a platform
that handles the downstream workflow.
CrocOTT covers the catalogue side:
keeping .vtt caption files synced with a VOD library and serving them at scale,
so the subtitle files produced during transcription plug directly into the content platform without
a manual upload step for each recording.
Frequently asked questions
Which free transcription tool has no file import limit?
WhoSayAh has no file import limit on the free tier. You can upload as many files as you need, subject to the per-file constraints of 100 MB and 60 minutes. No account is required and there is no monthly usage counter. Otter.ai's free plan, by contrast, allows only 3 file imports for the lifetime of the account.
Do free transcription tools produce accurate speaker labels?
Accuracy depends primarily on recording quality. A clear recording with one speaker at a time and minimal background noise produces reliable speaker attribution across all the tools in this list. Crosstalk - two voices overlapping - is the most common cause of mis-attribution. Entering the exact speaker count when the tool asks for it consistently improves results compared with using auto-detect, because auto-detect must estimate a number from the audio signal alone.
What export formats do free transcription tools support?
WhoSayAh exports .txt (plain text with speaker names prefixed to each line) and
.vtt (WebVTT with precise timestamps, ready to attach to a video as a subtitle
track). Otter.ai exports .txt, .docx and .pdf on the free
plan. Fireflies.ai and Notta restrict most export formats to paid plans; on the free tier,
transcripts are primarily readable on-screen within their apps.
Is the uploaded audio stored permanently?
Not on WhoSayAh. The uploaded file is deleted as soon as the transcript is ready, and the transcript is deleted after 24 hours. The privacy policy states this exactly - it is a retention commitment, not a marketing line. Other tools differ: Otter.ai stores past transcripts indefinitely as part of its searchable archive, which requires an account to access and persist.