What the work involves
- Data curation — collecting and cleaning domain audio and transcripts, handling licensing and PII, building a held-out test set that isn't leaking.
- Fine-tuning — full fine-tuning or parameter-efficient methods (LoRA / PEFT) with Hugging Face Transformers; managing catastrophic forgetting of general speech.
- Evaluation — WER and entity-level accuracy (drug names, statutes, tickers) on domain test sets, plus checks for hallucination on silence and music.
- Prompting & decoding — initial-prompt conditioning, vocabulary biasing, and post-processing for formatting and casing.
- Deployment — serving the fine-tuned model (often via faster-whisper / CTranslate2), routing between base and specialized models, monitoring WER in production.
Skills teams screen for
- Python and PyTorch; Hugging Face Transformers and Datasets
- Parameter-efficient fine-tuning (LoRA, QLoRA) and training-run hygiene (W&B / MLflow)
- ASR evaluation: WER/CER, alignment, test-set design, statistical significance
- Inference optimization: faster-whisper, CTranslate2, ONNX, quantization, batching
- Domain literacy for the target vertical (clinical, legal, finance) or the target language
Where these roles are
Two places: companies whose product depends on domain ASR, and any team running Whisper in production that has hit an accuracy ceiling.
- Healthcare — Abridge, Suki, DeepScribe, and Microsoft/Nuance build ambient clinical documentation. See medical ASR jobs.
- Legal — Verbit and Rev for court and deposition transcription. See legal transcription jobs.
- Speech platforms — AssemblyAI, Deepgram and Speechmatics fine-tune and train ASR models as their core work.
- In-house — media, education, contact-center and multilingual products that adopted Whisper and now need domain accuracy.
Compensation
These are ASR / speech-ML roles; pay tracks the wider market rather than a "fine-tuning" premium. For ranges and how to research a specific offer, see the Whisper jobs salary guide and the ASR engineer salary guide.