Whisper jobs in 2026

Since OpenAI released Whisper in 2022 it has become a default choice for speech recognition at startups and scale-ups. Companies increasingly hire for Whisper specifically — not just general ASR — to fine-tune, optimize, and run it in production. This page covers what those roles pay, the skills they screen for, and where they're posted.

Live openings: 4 Whisper roles from the last 2 digests — see all open roles.

Speech Science Technology ManagerPosted Oct 6
Motorola Solutions (Theatro) · Richardson, TX · $130K–$178K + incentive bonus
Leads the speech team behind Theatro, Motorola's voice-driven communication platform for retail store staff. The role architects the real-time pipeline (voice-command detection, VAD, ASR, TTS, noise suppression, echo cancellation, keyword spotting) and tunes ASR engines such as Whisper, Deepgram, Cerence and Sensory for a fixed command set. It also works with hardware teams on microphone arrays. Requires 7+ years in speech with at least 2 in management. The posting has been up 30+ days, so apply soon.
View posting →
Senior Software Engineer, Voice AIPosted Oct 6
Cloaked · New York, NY (hybrid) · $200K–$230K + equity + bonus
Works on Call Guard, an AI agent that answers unknown callers, holds a live conversation and blocks scams. It has screened 50M+ calls so far. You'd cut end-to-end STT → LLM → TTS latency and build defenses against voice-cloned attackers. Wants 5+ years and production experience with Whisper or Deepgram, ElevenLabs or Cartesia, and LiveKit or Pipecat. Telephony and SIP experience is a plus.
View posting →
Principal Speech Data LinguistPosted Sep 29
Innodata · Remote (US) · $160K–$185K
A rare one: owns transcription and segmentation standards end to end — verbatim and phonetic (IPA) conventions, timestamping, speaker labeling, code-switching — plus error taxonomies, inter-annotator agreement and the human-in-the-loop design question of where human expertise still beats ASR. Typically 8+ years in speech-data quality, a linguistics or phonetics degree, genuine IPA fluency, and working knowledge of Whisper, forced alignment and ELAN.
View posting →
Senior AI Engineer, Voice PlatformPosted Sep 29
ClickUp · Remote (US) · $200K–$250K
Owns the AI systems behind ClickUp's voice platform: streaming ASR, VAD and audio processing, accuracy gains through context injection (user names, teams, custom vocabulary, language detection), LLM post-processing for filler removal and formatting, and voice-to-action parsing into structured workspace commands. Also benchmarks and integrates third-party ASR (Whisper, AssemblyAI, Fireworks) on cost, latency and accuracy.
View posting →
Typical comp
$115K–$220K
Most-screened skill
Fine-tuning
Work style
Remote-heavy

Where Whisper roles are posted

There's no single board for these. In practice they show up on the career pages of companies doing transcription, meeting AI, and clinical documentation (see the list further down), on Greenhouse/Ashby/Lever, and in the weekly digest below. Search terms that surface them: "Whisper engineer", "speech recognition engineer", "ASR engineer", "audio ML engineer".

Current market pulse

Hiring Demand

Very High. Whisper has effectively become the "default" ASR choice for new products in 2026. Its combination of ease-of-use, multilingual support, and strong out-of-box accuracy makes it the obvious starting point for most companies. This creates consistent demand for engineers who can go beyond the basics to production-grade deployments.

Why companies want Whisper specialists:

Top Skills

Deep understanding of Whisper architecture, fine-tuning workflows with Hugging Face, and optimization techniques like Faster-Whisper and CTranslate2. Specific expertise in demand:

Compensation

Strong compensation driven by market demand. $140K-$200K total compensation is typical, with early-stage startups offering meaningful equity (0.2-0.8%) for engineers who can get their ASR system production-ready quickly.

Breakdown:

Common Use Cases You'll Build

Technical Challenges You'll Solve

Speed/Cost Optimization:

Accuracy Improvement:

Production Reliability:

Fine-Tuning Whisper: The Skill That Pays

Generic Whisper is good, but fine-tuned Whisper is great. Companies will pay premium for engineers who can:

Real results: Fine-tuning Whisper on 10-50 hours of domain-specific audio can reduce WER by 20-40% for that domain.

Companies Specifically Hiring Whisper Experts

Why Whisper Over Other ASR Systems?

Startups choose Whisper because:

Recommended Tools for Whisper Engineers

Note: Some of the links below are affiliate links. We may earn a small commission if you make a purchase through these links at no additional cost to you.

Hugging Face Audio Course

Free course specifically covering Whisper fine-tuning - essential learning

Start Free

Speech and Language Processing (Jurafsky)

Free online textbook - understand fundamentals beyond just using Whisper

Read Free

NVIDIA RTX 3060 (12GB)

Best budget GPU for Whisper development - enough VRAM for large-v3

View Options

Get the weekly Speech AI jobs digest

New ASR, TTS, voice-AI and speech-analytics roles from ~30 companies, plus salary and hiring notes — one email a week. Free, unsubscribe anytime.

✓ One email a week
✓ No spam, ever
✓ Unsubscribe anytime