"Speech analytics" covers three overlapping markets: revenue / conversation intelligence for sales teams, contact-center analytics and agent assist, and meeting intelligence. All of them run on the same core stack — ASR, speaker diarization, PII redaction, and then sentiment, topic and summarization models on top of the transcript, increasingly with an LLM in the loop.
For engineers, the roles are "speech + NLP + production ML." US total compensation for mid-to-senior roles broadly runs $150K–$260K; see our salary guide and interview questions. Live openings from most of these companies go into the weekly digest.
This is a field guide, not a ranking — grouped by what the company actually does.
Revenue & conversation intelligence
- Gong — analyzes sales calls, meetings and emails to surface deal risk and coaching signals. Large private company; hires ML/NLP engineers and applied scientists for transcription quality, diarization, and downstream deal-intelligence models.
- Chorus (ZoomInfo) — conversation intelligence for sales, acquired by ZoomInfo in 2021 and now part of its go-to-market platform. Public-company setting; speech and NLP roles within ZoomInfo's engineering org.
- Sybill — a newer entrant doing behavioral/emotion signals on sales calls; small ML team.
Contact-center analytics & agent assist
- Observe.AI — real-time agent assist, QA automation and coaching for contact centers. Real-time ASR plus LLM summarization; roles in streaming speech and ML infra.
- CallMiner — a long-established interaction-analytics vendor for enterprise and Fortune 500 contact centers, operating at very large call volumes. Roles span ASR, analytics pipelines, and on-prem/cloud deployment.
- Cresta — generative AI for contact centers (real-time guidance, knowledge assist). Speech + LLM roles.
- NICE and Verint — the incumbent enterprise CX / workforce-engagement platforms; large, established, lots of on-prem and regulated-industry work.
- Talkdesk — cloud contact center with an "Autopilot" voice-agent and analytics layer.
Meeting intelligence
- Otter.ai — meeting transcription, summaries and action items at consumer and business scale; senior applied-scientist roles owning ASR and TTS end to end.
- Fireflies.ai — an AI notetaker that joins calls across Zoom/Meet/Teams; remote-first, small team.
- Fathom — free meeting recorder and summarizer; earlier-stage.
Communications platforms with voice AI
- Dialpad — an AI business-phone / UCaaS platform whose "Ai" layer does real-time transcription, sentiment and coaching during calls. Real-time ASR + NLU roles.
- Aircall — SMB-focused cloud phone system, adding transcription and analytics features.
Speech APIs that power a lot of the above
Many analytics products build on a third-party ASR engine, and these companies hire for the models themselves. They also appear in our top companies hiring ASR engineers guide.
- AssemblyAI — speech-to-text plus audio-intelligence features (summaries, topics, redaction); research-driven.
- Deepgram — real-time / streaming ASR with its own models; strong on low latency.
- Speechmatics — UK enterprise ASR, wide language coverage, on-prem for regulated industries; UK visa sponsorship common.
What these roles involve
- ASR quality for messy real-world audio — telephony (8 kHz), crosstalk, accents, jargon.
- Speaker diarization — separating agent from customer, or every participant in a meeting; measured by DER.
- Redaction / compliance — removing PII and payment data from transcripts, often a hard requirement.
- Downstream models — sentiment, topics, call scoring, summarization; increasingly an LLM prompted over the transcript.
- Real-time systems for agent assist — latency budgets and reliability, which usually means an on-call rotation.
How to choose
- Enterprise vs. startup: incumbents (NICE, Verint, CallMiner) offer scale and stability with more process; newer companies offer more surface area and equity risk.
- Product type: B2B enterprise (contact center) vs. consumer/prosumer (meeting assistants) — different feedback loops and constraints.
- Batch vs. real-time: real-time agent assist pays a premium and carries on-call; post-call analytics is calmer.
- Stack: some teams run modern PyTorch/Whisper-era pipelines, others maintain older Kaldi-based systems — ask in the interview.
- Cash vs. equity: public or late-stage vs. early-stage.