Spoken NLP jobs in 2026

The convergence of speech and natural language processing has created one of the hottest specializations in 2026. As LLMs move from text-only to native audio processing (like GPT-4o and Gemini), companies are desperate for engineers who can bridge the gap between raw audio and semantic meaning—handling everything from intent extraction to conversational AI.

Live openings: 4 Spoken NLP roles from the last 2 digests — see all open roles.

Technical GTM Enablement ManagerPosted Oct 6
Deepgram · Remote (US) · compensation not listed
Deepgram's first hire in this role. You'd build the onboarding, playbooks and certifications that bring its sales engineers, solutions architects and customer engineers up to speed. That includes standards for technical discovery, proof-of-concept evaluations, architecture guidance and handoffs, plus readiness work before each product launch. Requires 5+ years in customer-facing technical roles and hands-on fluency with APIs and SDKs. You don't need to write production code, but "hand-waving does not pass." Experience with STT, TTS or voice agents is a plus. A good fit if you came up through solutions engineering and want a non-engineering role at a speech company.
View posting →
Senior Research Engineer, Audio and SpeechPosted Sep 29
Decagon · San Francisco or New York City (in-office) · $200K–$400K + equity
Builds the models and agent harnesses behind Decagon's real-time voice agents — full-duplex systems handling turn-taking, interruptions and overlapping speech, plus work on recognition, VAD, endpointing and generation across speakers and languages. Requires 4+ years in speech or multimodal ML and hands-on experience with autoregressive, diffusion, flow-matching or codec-based speech models. Their team publishes — worth reading their write-ups on flow-matching TTS with RL before you apply.
View posting →
Senior Applied Scientist, Real-Time Conversational AIPosted Sep 29
Amazon (AGI) · Sunnyvale, CA (onsite) · $192.2K–$260K + sign-on + RSUs
Large-scale multimodal foundation models for real-time speech and audio generation, including reward models that capture naturalness and conversational quality, and RL for natural timing. You own a research area across the full lifecycle, from pre-training and architecture through post-training alignment and real-time deployment. 5+ years in ML, or a PhD plus 6+ years; requires hands-on foundation-model training experience.
View posting →
Staff Software Engineer, Voice AgentPosted Sep 29
Decagon · San Francisco · $200K–$400K + equity
The architecture side of the same voice platform: owns Decagon's real-time voice runtime and its multi-quarter roadmap, sets reliability, testing and observability standards for live calls, and builds the frameworks that make voice systems debuggable. 8+ years with real technical leadership; speech recognition, VAD and streaming-protocol experience is listed as "even better if," not required — this is a real-time systems role first.
View posting →
Hiring Demand
Very High
Avg Salary
$180K-$250K
Growth Rate
+45% YoY

Current Market Pulse

Hiring Demand

Very High. The explosion of multimodal AI has created unprecedented demand for engineers who understand both speech recognition and natural language understanding. Voice assistants are evolving from simple command-response systems to full conversational agents that need to understand context, intent, emotion, and nuance from spoken input.

Major hiring sectors include:

Top Skills

Experience with Spoken Language Understanding (SLU), end-to-end audio-to-text-to-intent models, and NLP frameworks like Hugging Face or LangChain is essential. Specific skills in demand:

Compensation

Mid-to-senior roles are seeing explosive growth, often exceeding $180K–$250K total compensation at remote-first startups and FAANG companies. The scarcity of engineers who truly understand both domains (speech + NLP) commands a significant premium.

Salary breakdown:

Why This Niche is Exploding

The 2024-2026 shift from text-based LLMs to multimodal AI has fundamentally changed the landscape. Companies that built text-only NLP systems are now racing to add native audio understanding. This creates massive demand for "bridge" engineers who can:

Key Companies Hiring

Recommended Tools for Spoken NLP Engineers

Note: Some of the links below are affiliate links. We may earn a small commission if you make a purchase through these links at no additional cost to you.

Hugging Face Audio Course

Free comprehensive course covering speech + NLP integration - essential for this field

Start Free

Speech and Language Processing (Jurafsky & Martin)

The definitive textbook - free online version covers both speech and NLP fundamentals

Read Free

Blue Yeti USB Microphone

Professional audio quality for testing voice systems - under $100

View on Amazon

Get the weekly Speech AI jobs digest

New ASR, TTS, voice-AI and speech-analytics roles from ~30 companies, plus salary and hiring notes — one email a week. Free, unsubscribe anytime.

✓ One email a week
✓ No spam, ever
✓ Unsubscribe anytime