Most accuracy claims for speech to text come from clear, read speech. Real people are messier. They restart sentences, say “you know” and talk with an accent. So we took 55 minutes of real video call conversations from 24 people with American, British, Indian, Nigerian and Spanish accents, ran every recording through FreeTTS speech to text and six versions of OpenAI’s free Whisper, and counted every wrong, missing and extra word against the human transcripts.

Word error rate on the same 8,800 spoken words. Lower is better.
| Tool | American | British | Indian | Nigerian | Spanish accented | All 8,800 words |
|---|---|---|---|---|---|---|
| Whisper large-v3, fixed | 10.4% | 12.0% | 8.6% | 15.7% | 13.3% | 11.9% |
| Whisper large-v3-turbo, fixed | 11.6% | 11.6% | 9.2% | 14.9% | 14.1% | 12.2% |
| FreeTTS speech to text | 13.1% | 13.2% | 12.4% | 17.0% | 15.5% | 14.2% |
| Whisper large-v3 | 11.4% | 17.8% | 13.1% | 17.4% | 15.1% | 14.8% |
| Whisper medium | 11.9% | 14.2% | 13.9% | 17.7% | 16.6% | 14.8% |
| Whisper large-v3-turbo | 19.3% | 13.2% | 9.4% | 15.7% | 18.0% | 15.1% |
| Whisper small | 13.3% | 15.9% | 18.7% | 18.6% | 17.2% | 16.8% |
| Whisper base | 16.6% | 22.0% | 16.4% | 24.1% | 21.2% | 19.9% |
| Whisper tiny | 20.9% | 32.2% | 22.4% | 31.5% | 24.2% | 26.0% |
Word error rate = wrong, missing and extra words divided by the words actually spoken, pooled over every word in each group. Whisper always had the language set to English. “Fixed” means two settings changed: the voice activity filter on and “condition on previous text” off, explained below.
Everything here can be repeated: the recordings are public and every setting is listed.
We used EdAcc, the Edinburgh International Accents of English Corpus. It holds almost 40 hours of video call conversations between friends, transcribed word for word by professional transcribers, and a trained linguist labelled each speaker’s accent. The University of Edinburgh team published it for exactly this job, measuring speech recognition on accents, under a CC BY-SA 4.0 license.
From its test set we took five accent groups and about 10 minutes of speech from each, shared evenly between every speaker in the group, in the order they spoke:
| Accent group | Speakers | Minutes of speech | Words spoken |
|---|---|---|---|
| American | 9 | 12.8 | 1,888 |
| British | 2 | 10.4 | 1,580 |
| Indian | 3 | 10.2 | 1,933 |
| Nigerian | 5 | 10.9 | 1,745 |
| Spanish accented | 5 | 10.4 | 1,654 |
| Total | 24 | 54.6 | 8,800 |
The group names are EdAcc’s: Mainstream US English, Southern British English, Indian English, Nigerian English and Spanish. The Spanish group is English spoken by people whose first language is Spanish or Catalan. Some speakers in the American group grew up with another first language but were heard by the linguist as having a mainstream US accent.
We fixed the rules for picking clips before running any tool. We skipped clips the transcribers marked to be ignored, clips with two people talking at once or with words in another language, clips where people read out their participant codes, and clips shorter than one second or longer than 60 seconds. We removed sound labels such as laughter from the human transcripts. Each speaker’s clips were joined into one audio file with half a second of silence between them, and every tool got exactly the same files.
Word error rate is the standard measure: the words a tool got wrong, left out or added, divided by the words actually spoken. Before counting, both the tool’s text and the human transcript went through Whisper’s English text normaliser, which lowercases, removes punctuation, writes numbers and British spellings one way and drops “um”, “uh” and “hmm”. That way no tool loses points for “colour” against “color” or “25” against “twenty five”. We counted with the open source jiwer library and pooled the counts over every word in a group, so a long file counts more than a short one.
Default settings. The speaker counts are small, so read these as patterns, not rankings.
Whisper medium got 11.9% and FreeTTS 13.1%. Turbo’s 19.3% is partly one recording where it invented 68 words.
Whisper medium got 14.2%. Large-v3 scored 17.8%, partly because it got stuck repeating one sentence in one of the two recordings.
FreeTTS got 12.4% and large-v3 13.1%. Our three Indian speakers were among the easiest for most tools.
Whisper large-v3 got 17.4%, medium 17.7% and turbo 15.7%. One speaker in this group was hard for every tool.
FreeTTS got 15.5% and Whisper medium 16.6%. Turbo’s 18.0% includes a 36 word passage it made up.
Among the American speakers alone, FreeTTS got 1.5% of words wrong for one person and 29.3% for another. One Nigerian speaker came out between 33.1% and 55.9% on every tool we tried.
The EdAcc researchers saw the same kind of gap. In their paper, every model they tested did worse on Indian, Jamaican and Nigerian English speakers, and the best one got 2.7% of words wrong on clean US English but 19.7% on their conversations. Our Nigerian group fits that picture; our three Indian speakers did not, which is a good reminder that a handful of speakers can pull a group either way.
An error rate hides very different kinds of mistake. Here is what sat behind the numbers.
Every tool left words out, and most of them were small words: “I”, “is”, “you”, “it”. FreeTTS left out 765 of the 8,800, more than Whisper large-v3 (619) and turbo (590) but fewer than Whisper small (901). About one in five of the FreeTTS misses (152) were repeats such as “I I” or “the the”. It also tidies filler phrases. Our speakers said “you know” 55 times; FreeTTS left it out 32 times, Whisper medium 12 times and Whisper large-v3 5 times.
For captions, notes and articles that makes FreeTTS text easier to read. For a word for word record, such as research interviews or legal notes, it matters, and our scoring counts every one of those as an error.
The larger Whisper models added far more words that nobody said: 333 for large-v3-turbo and 265 for large-v3, against 104 for FreeTTS. Most are single words; a few are whole invented passages, which is the next section.
None of these tools knows the names of your places, products or people. Whatever you use, read names and numbers again before you publish.
It can invent text. In one American recording that every other tool got between 1.5% and 7.4% wrong, Whisper large-v3-turbo added 68 words that nobody said, including a Chinese character in the middle of an English sentence, and scored 36.1%. It also added a 36 word passage about photos to a Spanish accented recording. This is a known problem: a 2024 study by Koenecke and colleagues found that about 1% of Whisper transcriptions contained whole invented phrases or sentences.
It can get stuck in a loop.In one of our two British recordings, Whisper large-v3 wrote “but very soon after that I think we got cable” four times. That loop is part of the reason its British score was 17.8% while turbo managed 13.2%.
The fix.OpenAI’s own code notes that turning off “condition on previous text” makes the model “less prone to getting stuck in a failure loop, such as repetition looping”, and faster-whisper has a voice activity filter that skips stretches without speech. We ran medium, large-v3-turbo and large-v3 again with both settings:
| Whisper model | Default settings | With the fix | What changed |
|---|---|---|---|
| large-v3 | 14.8% | 11.9% | The loop is gone: that British recording fell from 18.1% to 9.2%. |
| large-v3-turbo | 15.1% | 12.2% | No invented text: the American recording fell from 36.1% to 3.0%, the Spanish accented one from 34.0% to 18.3%. |
| medium | 14.8% | 14.8% | No change overall. |
With the fix, large-v3 and turbo became the two most accurate setups in the whole test. The catch is that you have to install Whisper, pick the model and set these options yourself, and the large models need a strong computer. In faster-whisper the settings are vad_filter=True and condition_on_previous_text=False; OpenAI’s own command line tool has the second one as --condition_on_previous_text False.
In FreeTTS, Auto-detect is fine for English. It chose English (US) for all 24 speakers, and the result matched choosing English (US) yourself: 14.2% both ways. Picking the closest English did not help either. English (UK) gave our British speakers 13.1% against 13.2%, and English (India) gave our Indian speakers 13.4% against 12.4%.
In Whisper, set the language. We also asked each Whisper model to guess the language from the first 30 seconds of every file. Tiny, base and small all guessed Yoruba for one of our British speakers, and medium guessed Yoruba for one of our Nigerian speakers. Only large-v3 and large-v3-turbo got all 24 right. Whisper then transcribes the whole file in the language it guessed, so one wrong guess ruins the transcript.
Bigger models are slower. Speed is OpenAI’s own comparison against large; error rates are from our test.
| Model | Parameters | Speed vs large | Words wrong, default | With the fix |
|---|---|---|---|---|
| tiny | 39M | about 10x | 26.0% | not run |
| base | 74M | about 7x | 19.9% | not run |
| small | 244M | about 4x | 16.8% | not run |
| medium | 769M | about 2x | 14.8% | 14.8% |
| large-v3-turbo | 809M | about 8x | 15.1% | 12.2% |
| large-v3 | 1,550M | 1x | 14.8% | 11.9% |
Our pick if you run Whisper yourself: large-v3-turbo with the fix. It got 12.2%of words wrong, close to large-v3’s 11.9%, and OpenAI rates it about eight times faster. Tiny and base are quick but got 26.0% and 19.9% of words wrong.
On default settings it made the fewest mistakes of anything we tried: 14.2%. Leave the language on Auto-detect for English; picking English (UK) or English (India) did not help.
Turn the voice activity filter on and “condition on previous text” off. That took large-v3 from 14.8% to 11.9% and stopped the invented text and the loop.
Left to guess, tiny, base and small guessed Yoruba for one of our British speakers. Only large-v3 and turbo guessed right for all 24.
Holyrood Park became Hollywood Park in Whisper large-v3 and turbo, and Dubai became “the bite” in Whisper base. Names are the words readers notice first, so check them.
FreeTTS tidies fillers and repeats, which reads better but is not a word for word record. For interviews or legal notes, check what was left out.
If you want a transcript without installing anything, FreeTTS made the fewest mistakes of every tool on its default settings: 14.2% of words across five accents of real conversation. If you are happy to run Whisper yourself, large-v3 or large-v3-turbo with two settings changed did better still, at 11.9% and 12.2%.
Accent matters, but the speaker matters more: the same tool got one person almost perfect and another wrong on a third of the words. Whatever you use, read names and numbers again, and try it on a recording that sounds like yours before you commit.
Upload a recording or talk into your microphone. You can try FreeTTS speech to text without an account, and a free account covers 90 minutes a month.
In a September 2026 test on 8,800 words of real video call conversation from 24 speakers with American, British, Indian, Nigerian and Spanish accents (EdAcc corpus), OpenAI’s Whisper large-v3 got 11.9% of words wrong with its voice activity filter on and condition on previous text off, and 14.8% on default settings. FreeTTS speech to text got 14.2% on its defaults, the lowest of any setup without changes.
FreeTTS speech to text and six Whisper models on 24 EdAcc speakers, run on 27 September 2026 with the settings above.
Sanabria et al., The Edinburgh International Accents of English Corpus, ICASSP 2023; data on Hugging Face under CC BY-SA 4.0; EdAcc leaderboard.
openai/whisper (MIT license, model sizes and speed); Radford et al., 2022; English text normaliser.
SYSTRAN/faster-whisper, version 1.2.1, including the voice activity filter.
jiwer, the open source word error rate library.
Koenecke et al., Careless Whisper: Speech-to-Text Hallucination Harms, 2024.
Recordings and transcripts: EdAcc, University of Edinburgh, CC BY-SA 4.0. We did not copy or host any of the audio.