Hiring for speech recognition in 2026 is concentrated in four places: dedicated ASR-API companies, the big research labs, voice-agent startups that live or die on transcription quality, and vertical players in health, legal and automotive. Most roles are remote-friendly, and US total compensation typically runs from about $130K early-career to $300K+ at staff and research level (see our ASR engineer salary guide and Whisper salary guide).
This is a reference list, not a ranking — the "best" company depends on whether you want to train models, ship them to production, or embed them in a product. Live openings from most of these companies go into our weekly digest.
At a glance
| Company | What they build | Typical speech roles | Base |
|---|---|---|---|
| Deepgram | ASR + voice-agent API, own models | Speech research, inference, ML eng | US / remote |
| AssemblyAI | Speech-to-text + audio-intelligence API | Applied ML, speech research | SF / remote |
| Speechmatics | Multilingual ASR engine | Speech modeling, ML infra | Cambridge UK / remote |
| OpenAI | Whisper, realtime speech API | Research, applied speech | SF |
| Google / DeepMind | USM, Chirp, on-device ASR | Research, software eng | Global |
| Amazon | Alexa, AWS Transcribe | Applied scientist, SDE | Multiple |
| Meta AI (FAIR) | wav2vec, MMS, Seamless | Research scientist / engineer | Menlo Park / remote |
| NVIDIA | NeMo, Riva, Parakeet/Canary models | Speech LLM eng, research | Global |
| Otter.ai | Meeting transcription | Applied scientist, speech | Mountain View |
| Suki / Abridge / DeepScribe | Ambient clinical documentation | ASR / ML eng | US / remote |
Dedicated ASR / speech-to-text platforms
These companies sell transcription as the product, so nearly every ML, research and inference role touches speech directly.
- Deepgram — builds its own ASR and voice-agent models (the Nova family) and an API around them. Hires speech researchers, inference/optimization engineers, and ML engineers; strong on streaming and low latency.
- AssemblyAI — speech-to-text plus "audio intelligence" (summarization, topics, redaction). Applied-ML and research roles focused on model quality and new audio features.
- Speechmatics — a UK ASR company with a self-supervised engine covering a very wide language set. Speech-modeling and ML-infrastructure roles, often with UK visa sponsorship.
- Rev — large-scale transcription and captioning plus an API. Roles lean toward model operations, evaluation, and production ASR at volume.
- Gladia — a European ASR API company; smaller ML team, generalist speech/ML engineering.
- Otter.ai — meeting transcription and summarization at consumer scale; senior applied-scientist roles owning ASR and TTS end to end.
Big labs and platforms
The largest speech teams in the world. Expect a research track (publications at Interspeech / ICASSP / ASRU) and an engineering track shipping models into products used by hundreds of millions.
- OpenAI — Whisper and the realtime speech API; a small, senior speech team.
- Google / DeepMind — universal speech models (USM/Chirp), on-device recognition for Pixel and Android, and speech research across DeepMind.
- Amazon — Alexa and AWS Transcribe; one of the biggest ASR orgs, with applied-scientist and SDE roles across many sites.
- Apple — on-device speech recognition for Siri and dictation; privacy-constrained, efficiency-focused work.
- Microsoft — Azure AI Speech and the Nuance clinical/enterprise ASR business.
- Meta AI (FAIR) — wav2vec 2.0, MMS, and Seamless; foundational speech and speech-translation research.
- NVIDIA — the NeMo toolkit, Riva deployment stack, and the Parakeet/Canary model families; also speech-LLM and voice-agent roles.
Voice-agent and conversational-AI companies
ASR is upstream of everything these companies do, so speech-quality and latency work is first-class even when the job title says "ML engineer."
- Cartesia and Rime — TTS-led, but both hire for full-duplex and streaming speech pipelines.
- PolyAI, Vapi, Retell AI — voice agents for customer service; roles tuning and integrating ASR for telephony audio.
- Cresta, Observe.AI, Gong — contact-center and revenue intelligence; ASR plus diarization and analytics. (See also our speech analytics companies guide.)
Vertical ASR: health, legal, automotive
- Healthcare — Suki, Abridge, DeepScribe, and Microsoft/Nuance build ambient clinical documentation. The hard part is domain accuracy (drug names, abbreviations) and HIPAA-constrained pipelines. See medical ASR jobs.
- Legal — Verbit and Rev handle court and deposition transcription. See legal transcription jobs.
- Automotive — Cerence, SoundHound, and the automakers build in-car voice assistants with tight latency and noise constraints. See automotive voice AI jobs.
How to get hired
- Pick a track. Research (novel models, publications), applied ML (fine-tuning, evaluation, data), or systems/inference (latency, cost, serving). The interview loops are different.
- Ship something with real audio. A fine-tuned Whisper on a domain dataset, a streaming transcription demo, or a diarization + ASR pipeline with measured WER and DER beats a generic ML portfolio.
- Know the metrics. WER, CER, RTF (real-time factor), DER, and how to read an evaluation that isn't leaking.
- Contribute where the field lives. Whisper, faster-whisper, NeMo, icefall, SpeechBrain, pyannote — a merged PR gets noticed.
- Show up. Interspeech and ICASSP are the recruiting venues; for 2026 that also means SLT 2026 in Palermo.
Companies whose whole product is speech (Deepgram, AssemblyAI, Speechmatics) will often take a strong general ML engineer who can show they've gone deep on one audio project. Read how to break into speech tech.