Gradium's on-device, CPU-only text-to-speech for private voice AI - Voice AI Space Barcelona
Discover how Gradium implements private, high-quality text-to-speech using only CPU power, enabling secure, efficient, and accessible on-device voice AI solutions.
Summary
About Gradium
Gradium is a voice foundation model lab and a spin-off of Kyutai. The company trains its own models to perform text-to-speech, speech-to-text, and other audio-related tasks.
Key Elements of Voice AI
A successful voice AI system requires a balance of quality and scalability:
- Quality: Defined by a natural flow that sounds human rather than robotic, expressivity (including emotional control), robustness when handling complex inputs like email addresses and phone numbers, and full-duplex capabilities that allow mutual interruption.
- Scalability: Involves managing inference location and latency, controlling costs (especially given GPU scarcity), and ensuring privacy by keeping data local rather than sending it to the cloud.
Gradium Phonon
To address scalability and privacy, Gradium developed Phonon, an on-device text-to-speech model designed to run directly on a device's CPU (including smartphones, tablets, and laptops) without requiring a server connection. Key features of the model include:
- Resource Efficiency: The model has approximately 100 million parameters, an application size of 100 to 200 megabytes, and a memory footprint of less than 500 megabytes.
- Offline Functionality: Because it runs locally, the model can operate entirely in airplane mode.
- Low Latency: Phonon achieves a latency of approximately 30 to 35 milliseconds on a standard laptop and around 210 milliseconds on a Raspberry Pi.
- Voice Personalization: Users can clone a voice using a 10-second audio sample directly on the device.
- Language Support: The model currently supports English, French, German, Spanish, and Portuguese, with more languages in development. It achieves a word error rate of 0.83% without voice cloning and 1% with voice cloning.
Use Cases and Business Model
On-device voice AI is highly applicable for mobile games, consumer apps, and language learning platforms where cloud API costs would otherwise be prohibitive for free or low-cost users. Gradium offers Phonon under a license-based pricing model, charging a flat rate per user per month for unlimited usage.
Future Directions
Gradium is actively working on on-device speech-to-text (ASR) models to enable fully voice-driven interactive experiences. Additionally, the company is developing on-device live translation capabilities with zero-shot voice cloning, allowing real-time translation while preserving the speaker's original voice offline.
Related Content

Demo - ChickyTutor, an AI language tutor for everyone - Voice AI Space Amsterdam

Top Doctors' AI medical scribe for clinical reports - Voice AI Space Barcelona

Agora's real-time network powers conversational AI and avatars - Voice AI Space Barcelona

Enera's Voice AI agent for EV charger support - Voice AI Space Barcelona

Palabra AI launches fast real-time translation and TTS - Voice AI Space Barcelona

Detecting sarcasm with Voice AI - Voice AI Space Amsterdam

Vibe Coding a Voice AI Agent with Claude & Vapi - Voice AI Space Amsterdam

Lessons learnt from building an AI Voice Assistant for scientific labs - Voice AI Space Amsterdam