Principal Research Scientist - Speech and Audio Foundation Models
We are seeking a Principal Research Scientist to lead the development of speech and audio foundation models. The successful candidate will be responsible for building and improving voice and speech models, designing evaluation frameworks, and taking models into production. The ideal candidate will have hands-on experience with foundation model training and real voice or speech research.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
Principal Research Scientist Speech & Audio Foundation Models
Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.
$270,000โ$500,000 base plus bonus, equity and benefits (US).
Relocation assistance available. Visa transfer supported
San Francisco on-site preferred | Remote considered in the US, UK and parts of Europe.
Permanent, full-time.
A top end research lab building realtime voice models text-to-speech, speech-to-text and speech-to-speech delivered as an API. The models run in production behind consumer applications used at very large scale, across health, learning, therapy, companionship, media and gaming.
Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.
The role
Looking to advance your Development & Programming career with relocation support? Explore Development & Programming Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
Build the models that are the product! This is full-stack research ownership: you frame the question, run the experiments, and ship the result. Research is only finished when it is in production and measurable.
Responsibilities
- Train foundation models: pre-training, reinforcement learning, reward modelling, post-training, new architectures, scaling.
- Build and improve voice and speech models across TTS, STT and speech-to-speech.
- Design the evaluation that proves the work: benchmarks, eval loops, LLM-as-judge, failure analysis. Evaluation is treated as a research product in its own right, not as a pre-launch checkbox.
- Work on frontier problems adjacent to the roadmap: multimodal, agents and tool use, test-time compute.
- Take models into production alongside the serving engineering team, inside a sub-200ms latency budget and across 100+ languages.
Essential
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
- Hands-on foundation-model training. Pre-training, RL, reward modelling, post-training, scaling. Fine-tuning or building on top of someone else's model is a different discipline and is not what this role is.
- Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest signal; TTS and ASR both count. Text-only research does not transfer.
- Evidence you can point at: papers, shipped models, open-source contributions, or systems in production.
Desirable
- Evaluation depth: benchmarks, eval loops, quality measurement, failure analysis.
- Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech or ICASSP.
- PhD in ML or NLP, or equivalent practical experience you can point to.
- Frontier exposure: multimodal, agents, tool use, test-time compute.
- Public work: side projects, open-source, technical write-ups.
Similar Jobs
Explore other opportunities that match your interests