ASR Benchmarks 2026

Comprehensive word error rate (WER) comparison for modern speech recognition models

Last updated: January 18, 2026

This page provides up-to-date WER benchmarks for the most popular ASR models in 2026. All numbers are sourced from published papers, model cards, and verified community benchmarks. Use this as a reference when evaluating models for research or production deployment.

How to Read These Tables

Lower WER is better. WER measures the percentage of words incorrectly transcribed. A WER of 2.0% means 98% of words are correct. Production systems typically target <5% WER for good user experience.

New ASR research roles land every Tuesday at 6am PT. Free, one email a week.

LibriSpeech Benchmarks

LibriSpeech is the standard academic benchmark for English ASR. It consists of clean audiobook recordings. Most research papers report results on test-clean and test-other subsets.

Model test-clean test-other Parameters Year Tags
Whisper Large-v3 1.4% 2.9% 1550M 2023 Production
Conformer-CTC (Google) 1.9% 3.9% 600M 2020 SOTA 2020
Wav2Vec 2.0 Large 1.9% 3.5% 317M 2020 Self-Supervised
HuBERT Large 1.9% 3.3% 317M 2021 Self-Supervised
Whisper Medium 2.4% 4.9% 769M 2022 Production
ContextNet (Google) 2.1% 4.6% 112M 2020 Streaming
Kaldi Chain (TDNN-F) 3.2% 7.6% ~20M 2019 Production
Whisper Small 3.0% 5.8% 244M 2022 Edge-Friendly
Whisper Base 4.8% 8.3% 74M 2022 Edge-Friendly

Common Voice Benchmarks

Common Voice represents more realistic, diverse speech from crowd-sourced recordings. Higher WER is expected due to accent variation and recording quality.

Model English (test) Spanish (test) German (test) Notes
Whisper Large-v3 5.2% 6.8% 7.1% Multilingual training
Wav2Vec 2.0 XLSR-53 7.3% 9.2% 8.6% 53 languages
Whisper Medium 6.8% 8.5% 9.2% Cost-effective choice
Kaldi (monolingual) 12.4% 15.8% 14.2% Requires language-specific tuning

TED-LIUM 3 Benchmarks

TED talks represent challenging spontaneous speech with varied topics, accents, and speaking styles.

Model WER (test) Real-Time Factor Hardware
Whisper Large-v3 3.8% 0.15x A100 GPU
Conformer-Transducer 4.2% 0.22x V100 GPU
Wav2Vec 2.0 Large 4.8% 0.18x V100 GPU
Kaldi Chain 6.3% 0.08x CPU (16 cores)
Real-Time Factor Explained

RTF measures inference speed. 0.15x means processing 1 hour of audio takes 9 minutes. Lower is faster. Production systems typically need <0.3x for good UX.

Multilingual Performance

Cross-lingual ASR is critical for global products. These benchmarks show performance on 10 common languages.

Model Languages Avg WER Best For
Whisper Large-v3 99 7.2% General-purpose multilingual
Wav2Vec 2.0 XLSR-128 128 9.8% Low-resource languages
MMS (Meta) 1100+ 10.4% Rare language coverage
Google USM 100+ 8.6% Production streaming

Production Deployment Considerations

Benchmark WER doesn't always translate to production performance. Here's what matters for real-world deployments:

Factors Beyond WER

2026 Model Recommendations by Use Case

Use Case Recommended Model Why
Meeting Transcription Whisper Large-v3 Best accuracy, handles accents/noise
Real-Time Subtitles Conformer-RNN-T Streaming with low latency
Medical Transcription Fine-tuned Whisper Medium Cost-effective + customizable
Call Center Analytics Kaldi Chain CPU-efficient, proven at scale
On-Device (Mobile) Whisper Tiny + Quantization Small footprint, offline
Low-Resource Languages MMS or XLSR-128 1100+ language coverage

Working on ASR Research?

Labs and companies hiring researchers who understand these benchmarks post here. New roles every Tuesday at 6am PT.

Methodology & Sources

All benchmark numbers are sourced from:

Reporting Issues: If you find outdated or incorrect WER numbers, please email benchmarks@speechtechjobs.com with citations.

Related Resources