Open roles right now

Pulled from our last two weekly digests. Each role links straight to the company's own posting.

Principal Research & Engineering, Realtime Voice AIPosted Oct 6
Inflection AI · Palo Alto, CA · $400K–$550K base + equity
A hands-on lead who sets the technical roadmap for Inflection's real-time voice stack, the voice side of Pi. That covers streaming ASR, TTS, speech-to-speech, speech LLMs, turn-taking, barge-in and latency, plus build-vs-buy-vs-train decisions, with a 1,000-GPU cluster for experiments. The posting asks for evaluation that goes past WER, scoring interruption handling, emotional fit and task success. You'd also mentor a team while staying close to the code.
View posting →
Speech Science Technology ManagerPosted Oct 6
Motorola Solutions (Theatro) · Richardson, TX · $130K–$178K + incentive bonus
Leads the speech team behind Theatro, Motorola's voice-driven communication platform for retail store staff. The role architects the real-time pipeline (voice-command detection, VAD, ASR, TTS, noise suppression, echo cancellation, keyword spotting) and tunes ASR engines such as Whisper, Deepgram, Cerence and Sensory for a fixed command set. It also works with hardware teams on microphone arrays. Requires 7+ years in speech with at least 2 in management. The posting has been up 30+ days, so apply soon.
View posting →
Research Scientist II, Speech AI LabPosted Oct 6
Adobe Research · San Francisco, CA · $187.1K–$270.95K in California ($142.7K–$270.95K US range) + bonus
A senior audio research scientist for Adobe's speech generative AI and multimodal work, covering speech and audio generation, audio representations and large-scale training. The lab expects independent research leadership: first-author papers, patents and prototypes that product teams can ship. PhD preferred (a research-focused master's is accepted), and 3+ years of industry research is strongly preferred.
View posting →
Research Engineer, Language — Wearables Polyglot AIPosted Oct 6
Meta · Redmond, WA + 3 other locations · $154K–$217K + bonus + equity
Voice LLMs for speech recognition, translation and synthesis on Reality Labs wearables, deployed both on servers and on the device itself. Datasets and evaluation frameworks are part of the job, not an afterthought. Requires 5+ years building speech or language models and shipping them to production. Multilingual and edge-deployment experience are listed as pluses.
View posting →
Research EngineerPosted Oct 6
Sesame · San Francisco, Bellevue or New York (onsite) · $190K–$320K + stock options
An evaluation-first role on the team building Sesame's voice companion. You'd own the offline and live eval pipelines for its speech and multimodal models, build the dataset-curation tooling and monitoring, and scale training and inference for LLM-sized workloads. Requires expert-level PyTorch and evaluation metrics that "actually predict user happiness." This is a different opening from the Research Scientist role we listed on September 8.
View posting →
Audio ML Engineer (Research)Posted Oct 6
HARMAN · Northridge, CA (hybrid) · $134.25K–$196.9K
Perception models for HARMAN's Intelligent Audio research group: quality prediction, artifact detection, acoustic scene classification and listener-preference modeling. The models have to fit embedded and cloud budgets, using quantization, pruning and distillation where needed. Asks for 5+ years of applied ML, at least 2 of them on audio, speech or acoustics. The posting says shipped product impact counts for more than credentials.
View posting →
Senior Speech Software EngineerPosted Oct 6
ASAPP · New York or Mountain View (hybrid) · $215K–$235K + performance bonus
Half model tuning, half infrastructure. On the model side you'd adapt ASR and TTS for noisy call-center audio and improve TTS prosody and number pronunciation. On the systems side you'd build the multi-threaded server frameworks that run thousands of concurrent streaming ASR → LLM → TTS sessions. Requires 5+ years of distributed systems in Go or Python and hands-on ASR/TTS work. You'll need to know your codecs too: Opus, G.711, jitter and packet loss.
View posting →
Machine Learning Engineer, Real-Time Speech TranslationPosted Oct 6
LILT · Washington, D.C. or Boston ($129K–$161K), Indianapolis ($120K–$150K) · hybrid · US citizenship required
Owns the backend for LILT's new live-translation product. The ASR and MT models already exist; the job is wiring them into a low-latency streaming system on Ray Serve and GPU Kubernetes, including confidence scoring that sends uncertain segments to human linguists. Requires 3+ years of production Python with asyncio and hands-on WebSocket/gRPC streaming. The first languages are Japanese, Korean and English. US citizenship is a contract requirement.
View posting →
Senior Software Engineer, Voice AIPosted Oct 6
Cloaked · New York, NY (hybrid) · $200K–$230K + equity + bonus
Works on Call Guard, an AI agent that answers unknown callers, holds a live conversation and blocks scams. It has screened 50M+ calls so far. You'd cut end-to-end STT → LLM → TTS latency and build defenses against voice-cloned attackers. Wants 5+ years and production experience with Whisper or Deepgram, ElevenLabs or Cartesia, and LiveKit or Pipecat. Telephony and SIP experience is a plus.
View posting →
Product Engineer, Systems ArchitectPosted Oct 6
Wispr · San Francisco (onsite) · $220K–$300K (L5) or $270K–$350K (L6) + equity · visa sponsorship
Owns the detection, inference, state and delivery systems behind Wispr Flow, its system-wide dictation app, and Notetaker. A lot of the work is proving those systems are correct when no ground truth exists. The posting says speech experience is not required, and it names signal processing, ranking, fraud and ML evaluation as backgrounds that translate. That makes it one of the few ways into a voice company without a speech background.
View posting →
Technical GTM Enablement ManagerPosted Oct 6
Deepgram · Remote (US) · compensation not listed
Deepgram's first hire in this role. You'd build the onboarding, playbooks and certifications that bring its sales engineers, solutions architects and customer engineers up to speed. That includes standards for technical discovery, proof-of-concept evaluations, architecture guidance and handoffs, plus readiness work before each product launch. Requires 5+ years in customer-facing technical roles and hands-on fluency with APIs and SDKs. You don't need to write production code, but "hand-waving does not pass." Experience with STT, TTS or voice agents is a plus. A good fit if you came up through solutions engineering and want a non-engineering role at a speech company.
View posting →
Research Scientist Graduate, Seed Model — Speech (2027 start)Posted Oct 6
ByteDance · San Jose, CA · $218.4K–$387.6K base + bonus + RSUs
A new-grad role on the Seed Speech team, building speech foundation models for both understanding and generation: data construction, instruction tuning, alignment, and gains in recognition, synthesis and robustness. A BS is the minimum and an MS is preferred, with internship experience in speech or audio a plus. You can apply to at most two ByteDance roles worldwide and they're reviewed on a rolling basis, so apply early and put your graduation date on your resume.
View posting →
Software Engineer, Voice ModelPosted Sep 29
xAI · Palo Alto, CA · $150K–$450K + equity
Owns the Grok voice model pipeline end to end: speech data curation and synthetic generation, pre-training and post-training of speech-language models with SFT and RL, evaluation harnesses for accuracy, latency and expressiveness, then production integration. Wants a Python expert comfortable with JAX/PyTorch, Spark and Ray, and distributed training on Kubernetes. The range is unusually wide even by frontier-lab standards — assume level is negotiable.
View posting →
Senior Research Engineer, Audio and SpeechPosted Sep 29
Decagon · San Francisco or New York City (in-office) · $200K–$400K + equity
Builds the models and agent harnesses behind Decagon's real-time voice agents — full-duplex systems handling turn-taking, interruptions and overlapping speech, plus work on recognition, VAD, endpointing and generation across speakers and languages. Requires 4+ years in speech or multimodal ML and hands-on experience with autoregressive, diffusion, flow-matching or codec-based speech models. Their team publishes — worth reading their write-ups on flow-matching TTS with RL before you apply.
View posting →
Senior Applied Scientist, Real-Time Conversational AIPosted Sep 29
Amazon (AGI) · Sunnyvale, CA (onsite) · $192.2K–$260K + sign-on + RSUs
Large-scale multimodal foundation models for real-time speech and audio generation, including reward models that capture naturalness and conversational quality, and RL for natural timing. You own a research area across the full lifecycle, from pre-training and architecture through post-training alignment and real-time deployment. 5+ years in ML, or a PhD plus 6+ years; requires hands-on foundation-model training experience.
View posting →
Multimodal AI Researcher, AudioPosted Sep 29
Dolby · Atlanta, GA · $140.7K–$170K + bonus (equity for some roles)
Generative modeling for audio in Dolby's Advanced Technology Group: text-to-audio, video-to-audio and image-to-audio architectures, multimodal audio-video-text representations, source separation, speech enhancement and text-to-music, on the Machine Reasoning and Perception team. PhD required, along with a publication record at NeurIPS/ICLR/ICML or ICASSP/Interspeech. Research-track role — papers, IP and tech transfer to product groups.
View posting →
Research Scientist, Speech & AudioPosted Sep 29
Innodata · Remote (US) · $160K–$185K
Designs the data specifications and evaluation methodology that frontier labs buy — benchmarks for ASR, TTS, speech-to-speech, diarization and audio-language models, with a focus on robustness across accents, noise and code-switching, plus ablation studies proving which data decisions actually moved the model. Roughly 5+ years in speech or audio ML; hands-on with ESPnet, NeMo, SpeechBrain or Kaldi, and forced alignment.
View posting →
Principal Speech Data LinguistPosted Sep 29
Innodata · Remote (US) · $160K–$185K
A rare one: owns transcription and segmentation standards end to end — verbatim and phonetic (IPA) conventions, timestamping, speaker labeling, code-switching — plus error taxonomies, inter-annotator agreement and the human-in-the-loop design question of where human expertise still beats ASR. Typically 8+ years in speech-data quality, a linguistics or phonetics degree, genuine IPA fluency, and working knowledge of Whisper, forced alignment and ELAN.
View posting →
Audio Inference Engineer, Model EfficiencyPosted Sep 29
Cohere · Remote or New York, San Francisco, Toronto, Montreal · $205K–$380K (CA/NY/WA), $175K–$325K other US states, CA$295K–CA$535K Toronto/Montreal + equity
Pushes latency, throughput and quality on Cohere's audio model serving — profiling the system, finding bottlenecks, and building for streaming and duplex real-time workloads. C++ and Python, with GPU programming, multi-GPU parallelization and vLLM/SGLang/TensorRT-LLM experience as strong pluses. No minimum in-office requirement, but the team clusters in EST and PST.
View posting →
Staff Software Engineer, Voice AgentPosted Sep 29
Decagon · San Francisco · $200K–$400K + equity
The architecture side of the same voice platform: owns Decagon's real-time voice runtime and its multi-quarter roadmap, sets reliability, testing and observability standards for live calls, and builds the frameworks that make voice systems debuggable. 8+ years with real technical leadership; speech recognition, VAD and streaming-protocol experience is listed as "even better if," not required — this is a real-time systems role first.
View posting →

How this works

Rather than run a stale listings feed, we track the career pages of ~30 speech-AI companies and publish the real openings — usually 12–20 a week — in a free weekly digest. Each role links straight to the company's own posting.

Browse the specialty pages below for the market picture in each area (salary bands, in-demand skills, companies hiring), and subscribe to get the week's roles in your inbox.

Latest digest
Who's Hiring in Speech AI — October 6, 2026 · 12 roles
Read it

Browse the full archive →

Browse by specialty

Hiring?

If you're recruiting for a speech-tech role, tell us about it and we'll include it in the digest and on the relevant specialty page.