
Research Scientist
- Salary
- Not Mentioned
- Location
- Paris
- Arrangement
- On-site
- Posted
- today
Gradium is a frontier voice AI company on a mission to redefine how humans interact with machines.
Voice remains the most human, highest-stakes channel, but also the most broken. Long wait times, rigid IVRs, low automation, and poor handoffs create frustration on both sides of the line. Gradium is rebuilding voice from the ground up with proprietary models: real-time understanding, autonomous resolution, and seamless escalation when humans matter most.
We're a team of world-class talent, with initial traction and a clear belief: voice will be the next major frontier of applied AI. We recently raised a $100m seed round and are backed by top-tier investors. Our goal is not incremental improvement, it's to make AI-powered voice interactions feel reliable, scalable, and economically transformative.
Expect early-stage reality: high autonomy, fast decisions, unreasonable ambition, few handoffs, and very little process unless it earns its place.
The Role
We're looking for a Research Scientist to help build the next generation of voice AI models, across voice synthesis, transcription and recognition. You'll design and train the models our customers build on, and push the boundaries of what's possible in real-time voice.
You'll work at the core of the company, partnering closely with the founders, research, and product to take models from idea to production. This role exists because our edge is the quality and reliability of our models, and that starts with the science underneath. Expect to own hard problems end to end, from architecture to inference at scale.
What You'll Do
- Build state-of-the-art speech models: Design and implement models for voice synthesis and recognition that set the quality bar for the industry. You'll own architecture decisions and take models from research idea to something that holds up on real customer content.
- Make it fast enough for production: Optimize models for real-time inference at scale, so quality never comes at the cost of latency. Reason about the trade-offs between accuracy, speed, and cost, and get the most out of the hardware.
- Own training end to end: Build and maintain the training pipelines for large-scale model training, from data to distributed runs. Keep experiments fast and reproducible so the team can iterate quickly.
- Turn research into product: Stay on top of the latest advances in speech AI and bring the ones that matter into our models. Work closely with the product team to translate customer requirements into technical solutions that ship.
Who You Are
- Deep speech and AI expertise: You have an MS or PhD in Computer Science, Machine Learning, AI or a related field, and 3+ years in machine learning with a focus on speech or audio. You know deep learning architectures (Transformers, CNNs, RNNs) cold and understand how to make them work in practice.
- Strong engineer, not just a researcher: You write excellent Python and are fluent with PyTorch or TensorFlow. You've run large-scale distributed training and know what breaks at scale.
- Founder mindset: You act with urgency, take full ownership, and don't wait for permission or perfect information. You are comfortable making high-stakes decisions in ambiguous environments and see the founding team as partners, not hierarchy.
- Obsessed with real-world quality: You care about how a model behaves on messy, real customer content, not just benchmark scores. You've felt the gap between a research demo and a production system, and you close it.
Nice to have: published research at top-tier ML venues (NeurIPS, ICML, ICLR); hands-on experience with speech synthesis models (Tacotron, FastSpeech, VITS); knowledge of audio signal processing and acoustic modeling; or experience with model optimization and quantization.
Sourced from Gradium’s careers page — applying takes you there.