How is word error rate calculated?
WER = (substitutions + deletions + insertions) ÷ reference words. The comparison aligns the transcript with a reference and finds a minimum edit count. Replacing one word, omitting one word and adding one word each count as an error under this simple definition. The resulting rate can exceed 100% when insertions outweigh a short reference.
The NIST Speech Recognition Scoring Toolkit provides established recognition-scoring tools. Our small downloadable implementation is a teaching example, not a claim of compatibility with every NIST option or vendor evaluation. OpenAI's Whisper repository also reports performance across specific languages and datasets, illustrating why a model name alone does not define one universal accuracy number.
Inspect the seven executed examples
The exact-match example has zero errors. Changing “the” to “an” in “send the invoice today” gives one substitution over four reference words: 25%. Removing “not” from “do not send the invoice” gives one deletion over five words: 20%. Adding “please” and “now” to a three-word reference gives two insertions: about 66.7%.
Our other cases cover punctuation/case normalization, a 300% score on a one-word reference and an empty reference. Empty-reference WER is recorded as null, because dividing by zero would not create a meaningful percentage. All seven expected edit-count checks passed locally. Every text pair is invented; no audio or speech recognizer was involved.
python3 score-wer.py wer-fixtures.json- Inspect the teaching scorer · Python
- Seven invented text pairs · JSON
- Actual local scoring output · JSON
- Normalization and scope · Markdown
Why normalization changes the result
The scorer uses Unicode case folding, extracts word-like tokens and keeps internal ASCII apostrophes. It discards punctuation. “Hello, TEAM!” and “hello team” therefore produce the same tokens. A punctuation-sensitive measure would answer a different question. Numbers written as digits versus words, abbreviations, speaker labels and disfluencies also require a stated policy.
Equal-cost alignments can assign different substitution, deletion and insertion counts. This implementation favors substitution, then deletion, then insertion when total errors tie. Report that policy instead of assuming every scoring package will produce identical component counts.
Low WER does not settle usefulness
The deletion of “not” changes the invoice instruction's meaning even though it is only one word. A name, price or deadline can similarly matter more than a filler word. Review important fields separately and measure the correction effort needed for the final task. WER does not evaluate speaker attribution, timestamps, formatting, privacy or the quality of a downstream summary.
When comparing two systems, use the same permitted audio, reference and normalization. Record the exact model, language, settings and failures. Keep the test set fixed and disclose how it was chosen. A single short fixture cannot justify a universal winner or a percentage claim about all meetings.
Connect accuracy to the actual work
For a podcast, your acceptance checks may include names, segment timestamps and usable edit points. For a searchable internal recording, terminology and retrieval may matter more. Start with those requirements, then use WER as one inspectable measurement. The audio-cost worksheet keeps duration and review work separate from the raw service charge. The ElevenLabs/Whisper reference remains limited; these local scoring examples do not complete its head-to-head comparison.
Sources and scope
- NIST Speech Recognition Scoring Toolkit · Read 10 October 2026.
- OpenAI Whisper repository · Read 10 October 2026.
Factual explanation and local teaching examples. Source reads and local executions have separate scopes. No native provider workflow, send, payment, deliverability result or comparative winner is claimed. Send a correction.