All writingAI & product

Speech to text accuracy: the demo that lies to you

Vendor speech to text accuracy numbers are measured on clean audio and neutral accents. Your calls are neither. Here is how to test properly.

Muhammad AbueleninCo-Founder28 Sept 20265 min read
AI & productCallix

Speech to text accuracy is the share of spoken words a system transcribes correctly, and every vendor quotes a figure that is true under the conditions it was measured. Those conditions are almost never yours: a quiet room, a good microphone, one speaker at a time, and an accent the model saw plenty of during training. Test on your own audio and the number moves, usually by more than the gap between the vendors you are comparing.

Real calls are a phone line at 8kHz, a reception desk with a printer running, two people talking over each other, and a caller switching between languages mid-sentence. Accuracy on that audio is the only number that matters, and it is not the number on the slide. It is also the foundation everything in call intelligence is built on.

What does word error rate hide?

Word error rate counts every mistake equally. Missing 'the' costs the same as missing 'not'. For call intelligence this is badly misleading, because the words that carry meaning are precisely the ones that are hardest to hear: names, numbers, drug names, product names, prices, and negations.

A transcript at 95% word error rate that reliably drops the word 'not' is worse than useless. It will confidently tell you a customer agreed to something they refused, and every downstream score, every alert, and every coaching conversation inherits that error without anyone seeing where it came from.

The distinction that matters is between the two ways a system can be wrong. It can miss a word, which leaves a hole you can usually spot. Or it can invent one, replacing an unclear phrase with a confident guess. Invention is rarer, far more damaging, and almost never reported in a vendor benchmark.

Why speech to text accuracy collapses on accents

Models degrade unevenly across accents, and the degradation follows the training data rather than anything about the speaker. A system that handles one regional accent flawlessly may struggle with another from fifty miles away, because one of them was well represented in the corpus and the other was not.

Code-switching, a caller moving between two languages inside one sentence, breaks many systems entirely. Language is typically detected once at the start of a call and then assumed for its duration. Multilingual call bases are common in the operations we work with, and a vendor that transcribes both languages competently in isolation can still fail on the switch, which looks fine in a demo and fails on your floor.

This is the part of speech to text accuracy that a single headline figure cannot express, because it is not one number. It is a different number for every accent, line quality, and language mix in your call base, and the average hides the segment where your system is failing hardest.

Two things follow. First, an aggregate accuracy figure is only meaningful next to the call mix it was measured on. Second, if a meaningful share of your customers are non-native speakers, that segment deserves its own measurement rather than a footnote in someone else's average.

One accuracy figure, five different answers
Overall words correct against recall on the terms that matter
All words Prices, names, negations
Vendor benchmark96% / 94%
Clean VoIP, native93% / 88%
Landline, native88% / 79%
Landline, second language81% / 64%
Code-switching mid-call74% / 51%

The headline number is the top row. The row your scoring actually depends on is the bottom one, on the segment you serve most.

One headline accuracy figure broken into the segments it averages over, with recall on the terms your scoring actually depends on.

How do you test transcription on your own calls?

The whole test takes an afternoon and it is worth more than six weeks of comparing published benchmarks.

  1. 1Pull twenty real calls. Not clean ones: pick five great, ten typical, and five that were hard to hear.
  2. 2Have a person transcribe only the moments that matter: the price quoted, the objection, the commitment, the outcome. Not the whole call.
  3. 3Run the same twenty through the system and compare only those moments.
  4. 4Count separately: things missed, things invented, and speakers mislabelled. Invention is the most dangerous and the least reported.
  5. 5Repeat with the accents and languages that make up your real call mix, in roughly the proportions they occur.

Two details decide whether the test is honest. Use recordings from the same telephony path you will use in production, because a codec change alone can move the result. And have someone who was not involved in choosing the vendor do the human transcription, for the same reason you would not let a candidate mark their own exam.

What speech to text accuracy bar is good enough?

It is easy to get absorbed in transcription quality and forget it is an intermediate artefact. Nobody's job improves because a transcript is two percent better. The job improves when the right call reaches the right manager the same day.

So the bar is set by the decision sitting on top of the transcript, not by a leaderboard. If you are routing complaints to a supervisor, you need the complaint phrases caught reliably and can tolerate a scruffy transcript elsewhere. If you are scoring compliance disclosures word by word, the bar on that passage is much higher, and the rest of the call barely matters.

Set the bar high enough that the scoring on top of it is trustworthy, then spend every remaining hour of the evaluation on what the system does with the words. That is where the gap between vendors is far larger, and it is the part no one benchmarks. Scoring, routing, and next actions are what the platform is judged on once the words are right.

We spent six weeks comparing word error rates and about forty minutes asking what anyone would do with the output. That ratio was exactly backwards.
Head of Revenue Operations, insurance group

Where this advice stops applying

If your calls are recorded over a clean VoIP connection, in one language, with a customer base that shares the accent the model was trained on, published accuracy figures will be roughly right and this whole exercise is overhead. That describes very few contact centres, but it does describe some.

The advice also stops short of a guarantee. Testing twenty calls tells you whether a system is fit for your audio; it does not tell you how it will behave on a caller type you have never had. Re-run the same afternoon whenever your call mix changes materially, using your own recordings rather than a scripted demo, and treat any accuracy figure, including one you measured yourself, as a claim with an expiry date.

speech to text accuracycall transcription softwaremultilingual call transcriptionword error rate
Keep reading
OperationsCallix

Your first contact resolution rate is a survey artefact

Your first contact resolution rate comes from a fifth of your customers and is reported as if it came from all of them. The call data already knows better.

5 min read
OperationsCallix

Repeat calls are a routing problem

Repeat calls cluster by topic, not by agent. So coaching the people who handle those topics will not move the number, and your own data shows why.

5 min read

Put your sales floor
on autopilot.

See what Callix hears in your very next call. Book a demo.