Principal Research & Engineering, Realtime Voice AIPosted Oct 6
Inflection AI · Palo Alto, CA · $400K–$550K base + equity
A hands-on lead who sets the technical roadmap for Inflection's real-time voice stack, the voice side of Pi. That covers streaming ASR, TTS, speech-to-speech, speech LLMs, turn-taking, barge-in and latency, plus build-vs-buy-vs-train decisions, with a 1,000-GPU cluster for experiments. The posting asks for evaluation that goes past WER, scoring interruption handling, emotional fit and task success. You'd also mentor a team while staying close to the code.
View posting →
Speech Science Technology ManagerPosted Oct 6
Motorola Solutions (Theatro) · Richardson, TX · $130K–$178K + incentive bonus
Leads the speech team behind Theatro, Motorola's voice-driven communication platform for retail store staff. The role architects the real-time pipeline (voice-command detection, VAD, ASR, TTS, noise suppression, echo cancellation, keyword spotting) and tunes ASR engines such as Whisper, Deepgram, Cerence and Sensory for a fixed command set. It also works with hardware teams on microphone arrays. Requires 7+ years in speech with at least 2 in management. The posting has been up 30+ days, so apply soon.
View posting →
Research Scientist II, Speech AI LabPosted Oct 6
Adobe Research · San Francisco, CA · $187.1K–$270.95K in California ($142.7K–$270.95K US range) + bonus
A senior audio research scientist for Adobe's speech generative AI and multimodal work, covering speech and audio generation, audio representations and large-scale training. The lab expects independent research leadership: first-author papers, patents and prototypes that product teams can ship. PhD preferred (a research-focused master's is accepted), and 3+ years of industry research is strongly preferred.
View posting →
Research Engineer, Language — Wearables Polyglot AIPosted Oct 6
Meta · Redmond, WA + 3 other locations · $154K–$217K + bonus + equity
Voice LLMs for speech recognition, translation and synthesis on Reality Labs wearables, deployed both on servers and on the device itself. Datasets and evaluation frameworks are part of the job, not an afterthought. Requires 5+ years building speech or language models and shipping them to production. Multilingual and edge-deployment experience are listed as pluses.
View posting →
Research EngineerPosted Oct 6
Sesame · San Francisco, Bellevue or New York (onsite) · $190K–$320K + stock options
An evaluation-first role on the team building Sesame's voice companion. You'd own the offline and live eval pipelines for its speech and multimodal models, build the dataset-curation tooling and monitoring, and scale training and inference for LLM-sized workloads. Requires expert-level PyTorch and evaluation metrics that "actually predict user happiness." This is a different opening from the Research Scientist role we listed on September 8.
View posting →
Audio ML Engineer (Research)Posted Oct 6
HARMAN · Northridge, CA (hybrid) · $134.25K–$196.9K
Perception models for HARMAN's Intelligent Audio research group: quality prediction, artifact detection, acoustic scene classification and listener-preference modeling. The models have to fit embedded and cloud budgets, using quantization, pruning and distillation where needed. Asks for 5+ years of applied ML, at least 2 of them on audio, speech or acoustics. The posting says shipped product impact counts for more than credentials.
View posting →