Videoby Voice AI Space

    Gradium's on-device, CPU-only text-to-speech for private voice AI - Voice AI Space Barcelona

    Discover how Gradium implements private, high-quality text-to-speech using only CPU power, enabling secure, efficient, and accessible on-device voice AI solutions.

    Summary

    About Gradium

    Gradium is a voice foundation model lab and a spin-off of Kyutai. The company trains its own models to perform text-to-speech, speech-to-text, and other audio-related tasks.

    Key Elements of Voice AI

    A successful voice AI system requires a balance of quality and scalability:

    • Quality: Defined by a natural flow that sounds human rather than robotic, expressivity (including emotional control), robustness when handling complex inputs like email addresses and phone numbers, and full-duplex capabilities that allow mutual interruption.
    • Scalability: Involves managing inference location and latency, controlling costs (especially given GPU scarcity), and ensuring privacy by keeping data local rather than sending it to the cloud.

    Gradium Phonon

    To address scalability and privacy, Gradium developed Phonon, an on-device text-to-speech model designed to run directly on a device's CPU (including smartphones, tablets, and laptops) without requiring a server connection. Key features of the model include:

    • Resource Efficiency: The model has approximately 100 million parameters, an application size of 100 to 200 megabytes, and a memory footprint of less than 500 megabytes.
    • Offline Functionality: Because it runs locally, the model can operate entirely in airplane mode.
    • Low Latency: Phonon achieves a latency of approximately 30 to 35 milliseconds on a standard laptop and around 210 milliseconds on a Raspberry Pi.
    • Voice Personalization: Users can clone a voice using a 10-second audio sample directly on the device.
    • Language Support: The model currently supports English, French, German, Spanish, and Portuguese, with more languages in development. It achieves a word error rate of 0.83% without voice cloning and 1% with voice cloning.

    Use Cases and Business Model

    On-device voice AI is highly applicable for mobile games, consumer apps, and language learning platforms where cloud API costs would otherwise be prohibitive for free or low-cost users. Gradium offers Phonon under a license-based pricing model, charging a flat rate per user per month for unlimited usage.

    Future Directions

    Gradium is actively working on on-device speech-to-text (ASR) models to enable fully voice-driven interactive experiences. Additionally, the company is developing on-device live translation capabilities with zero-shot voice cloning, allowing real-time translation while preserving the speaker's original voice offline.