Speech to text accuracy is the share of spoken words a system transcribes correctly, and every vendor quotes a figure that is true under the conditions it was measured. Those conditions are almost never yours: a quiet room, a good microphone, one speaker at a time, and an accent the model saw plenty of during training. Test on your own audio and the number moves, usually by more than the gap between the vendors you are comparing.
Real calls are a phone line at 8kHz, a reception desk with a printer running, two people talking over each other, and a caller switching between languages mid-sentence. Accuracy on that audio is the only number that matters, and it is not the number on the slide. It is also the foundation everything in call intelligence is built on.
What does word error rate hide?
Word error rate counts every mistake equally. Missing 'the' costs the same as missing 'not'. For call intelligence this is badly misleading, because the words that carry meaning are precisely the ones that are hardest to hear: names, numbers, drug names, product names, prices, and negations.
A transcript at 95% word error rate that reliably drops the word 'not' is worse than useless. It will confidently tell you a customer agreed to something they refused, and every downstream score, every alert, and every coaching conversation inherits that error without anyone seeing where it came from.
The distinction that matters is between the two ways a system can be wrong. It can miss a word, which leaves a hole you can usually spot. Or it can invent one, replacing an unclear phrase with a confident guess. Invention is rarer, far more damaging, and almost never reported in a vendor benchmark.
Why speech to text accuracy collapses on accents
Models degrade unevenly across accents, and the degradation follows the training data rather than anything about the speaker. A system that handles one regional accent flawlessly may struggle with another from fifty miles away, because one of them was well represented in the corpus and the other was not.
Code-switching, a caller moving between two languages inside one sentence, breaks many systems entirely. Language is typically detected once at the start of a call and then assumed for its duration. Multilingual call bases are common in the operations we work with, and a vendor that transcribes both languages competently in isolation can still fail on the switch, which looks fine in a demo and fails on your floor.
This is the part of speech to text accuracy that a single headline figure cannot express, because it is not one number. It is a different number for every accent, line quality, and language mix in your call base, and the average hides the segment where your system is failing hardest.
Two things follow. First, an aggregate accuracy figure is only meaningful next to the call mix it was measured on. Second, if a meaningful share of your customers are non-native speakers, that segment deserves its own measurement rather than a footnote in someone else's average.
The headline number is the top row. The row your scoring actually depends on is the bottom one, on the segment you serve most.
How do you test transcription on your own calls?
The whole test takes an afternoon and it is worth more than six weeks of comparing published benchmarks.
- 1Pull twenty real calls. Not clean ones: pick five great, ten typical, and five that were hard to hear.
- 2Have a person transcribe only the moments that matter: the price quoted, the objection, the commitment, the outcome. Not the whole call.
- 3Run the same twenty through the system and compare only those moments.
- 4Count separately: things missed, things invented, and speakers mislabelled. Invention is the most dangerous and the least reported.
- 5Repeat with the accents and languages that make up your real call mix, in roughly the proportions they occur.
Two details decide whether the test is honest. Use recordings from the same telephony path you will use in production, because a codec change alone can move the result. And have someone who was not involved in choosing the vendor do the human transcription, for the same reason you would not let a candidate mark their own exam.
What speech to text accuracy bar is good enough?
It is easy to get absorbed in transcription quality and forget it is an intermediate artefact. Nobody's job improves because a transcript is two percent better. The job improves when the right call reaches the right manager the same day.
So the bar is set by the decision sitting on top of the transcript, not by a leaderboard. If you are routing complaints to a supervisor, you need the complaint phrases caught reliably and can tolerate a scruffy transcript elsewhere. If you are scoring compliance disclosures word by word, the bar on that passage is much higher, and the rest of the call barely matters.
Set the bar high enough that the scoring on top of it is trustworthy, then spend every remaining hour of the evaluation on what the system does with the words. That is where the gap between vendors is far larger, and it is the part no one benchmarks. Scoring, routing, and next actions are what the platform is judged on once the words are right.
We spent six weeks comparing word error rates and about forty minutes asking what anyone would do with the output. That ratio was exactly backwards.
Where this advice stops applying
If your calls are recorded over a clean VoIP connection, in one language, with a customer base that shares the accent the model was trained on, published accuracy figures will be roughly right and this whole exercise is overhead. That describes very few contact centres, but it does describe some.
The advice also stops short of a guarantee. Testing twenty calls tells you whether a system is fit for your audio; it does not tell you how it will behave on a caller type you have never had. Re-run the same afternoon whenever your call mix changes materially, using your own recordings rather than a scripted demo, and treat any accuracy figure, including one you measured yourself, as a claim with an expiry date.