Top companies hiring ASR & speech engineers in 2026

Hiring for speech recognition in 2026 is concentrated in four places: dedicated ASR-API companies, the big research labs, voice-agent startups that live or die on transcription quality, and vertical players in health, legal and automotive. Most roles are remote-friendly, and US total compensation typically runs from about $130K early-career to $300K+ at staff and research level (see our ASR engineer salary guide and Whisper salary guide).

This is a reference list, not a ranking — the "best" company depends on whether you want to train models, ship them to production, or embed them in a product. Live openings from most of these companies go into our weekly digest.

At a glance

CompanyWhat they buildTypical speech rolesBase
DeepgramASR + voice-agent API, own modelsSpeech research, inference, ML engUS / remote
AssemblyAISpeech-to-text + audio-intelligence APIApplied ML, speech researchSF / remote
SpeechmaticsMultilingual ASR engineSpeech modeling, ML infraCambridge UK / remote
OpenAIWhisper, realtime speech APIResearch, applied speechSF
Google / DeepMindUSM, Chirp, on-device ASRResearch, software engGlobal
AmazonAlexa, AWS TranscribeApplied scientist, SDEMultiple
Meta AI (FAIR)wav2vec, MMS, SeamlessResearch scientist / engineerMenlo Park / remote
NVIDIANeMo, Riva, Parakeet/Canary modelsSpeech LLM eng, researchGlobal
Otter.aiMeeting transcriptionApplied scientist, speechMountain View
Suki / Abridge / DeepScribeAmbient clinical documentationASR / ML engUS / remote

Dedicated ASR / speech-to-text platforms

These companies sell transcription as the product, so nearly every ML, research and inference role touches speech directly.

  • Deepgram — builds its own ASR and voice-agent models (the Nova family) and an API around them. Hires speech researchers, inference/optimization engineers, and ML engineers; strong on streaming and low latency.
  • AssemblyAI — speech-to-text plus "audio intelligence" (summarization, topics, redaction). Applied-ML and research roles focused on model quality and new audio features.
  • Speechmatics — a UK ASR company with a self-supervised engine covering a very wide language set. Speech-modeling and ML-infrastructure roles, often with UK visa sponsorship.
  • Rev — large-scale transcription and captioning plus an API. Roles lean toward model operations, evaluation, and production ASR at volume.
  • Gladia — a European ASR API company; smaller ML team, generalist speech/ML engineering.
  • Otter.ai — meeting transcription and summarization at consumer scale; senior applied-scientist roles owning ASR and TTS end to end.

Big labs and platforms

The largest speech teams in the world. Expect a research track (publications at Interspeech / ICASSP / ASRU) and an engineering track shipping models into products used by hundreds of millions.

  • OpenAI — Whisper and the realtime speech API; a small, senior speech team.
  • Google / DeepMind — universal speech models (USM/Chirp), on-device recognition for Pixel and Android, and speech research across DeepMind.
  • Amazon — Alexa and AWS Transcribe; one of the biggest ASR orgs, with applied-scientist and SDE roles across many sites.
  • Apple — on-device speech recognition for Siri and dictation; privacy-constrained, efficiency-focused work.
  • Microsoft — Azure AI Speech and the Nuance clinical/enterprise ASR business.
  • Meta AI (FAIR) — wav2vec 2.0, MMS, and Seamless; foundational speech and speech-translation research.
  • NVIDIA — the NeMo toolkit, Riva deployment stack, and the Parakeet/Canary model families; also speech-LLM and voice-agent roles.

Voice-agent and conversational-AI companies

ASR is upstream of everything these companies do, so speech-quality and latency work is first-class even when the job title says "ML engineer."

Vertical ASR: health, legal, automotive

How to get hired

  1. Pick a track. Research (novel models, publications), applied ML (fine-tuning, evaluation, data), or systems/inference (latency, cost, serving). The interview loops are different.
  2. Ship something with real audio. A fine-tuned Whisper on a domain dataset, a streaming transcription demo, or a diarization + ASR pipeline with measured WER and DER beats a generic ML portfolio.
  3. Know the metrics. WER, CER, RTF (real-time factor), DER, and how to read an evaluation that isn't leaking.
  4. Contribute where the field lives. Whisper, faster-whisper, NeMo, icefall, SpeechBrain, pyannote — a merged PR gets noticed.
  5. Show up. Interspeech and ICASSP are the recruiting venues; for 2026 that also means SLT 2026 in Palermo.
If you don't have speech experience yet

Companies whose whole product is speech (Deepgram, AssemblyAI, Speechmatics) will often take a strong general ML engineer who can show they've gone deep on one audio project. Read how to break into speech tech.

Get these companies' openings weekly

We track the careers pages above and put the real roles in one email a week — free.