Find all the Voice AI Startup programs:
    Videoby Voice AI Space

    How Speech-to-Speech AI Models Work: Audio Tokens to Full-Duplex

    Master speech to speech AI architecture by exploring audio tokenization and full duplex systems for seamless real time voice interaction.

    Summary

    Speech-to-Speech Models vs. Three-Model Systems

    Traditional voice agents operate like a relay race, using three separate models: speech-to-text to transcribe the input, a language model to generate a text response, and text-to-speech to read the answer aloud. In contrast, a speech-to-speech model processes everything within a single model, taking sound directly as input and producing sound as output without passing a text transcript in between.

    How Speech-to-Speech Models Work

    Just as text models predict the next token, speech-to-speech models predict the next short slice of sound. To do this, sound must first be converted into discrete pieces using an audio codec. The codec cuts the audio into short slices and describes each slice with numbers from fixed lists called codebooks. For example, Kyutai's Moshi model uses 12.5 slices per second with 8 numbers per slice, where the first number represents what was said and the remaining seven represent how it sounded (voice, tone, and detail).

    Hearing, Speaking, and the "Text Trick"

    Models process audio input in one of two ways: by reading audio tokens directly or by using an encoder (a listening network) to convert sound into numbers for the language model. Speaking reverses this process, using a decoder to turn predicted audio tokens back into sound. To maintain real-time performance, models like Moshi split the workload between a large model that runs once per slice and a smaller model that fills in the remaining audio numbers.

    Because audio tokens spread meaning thinly, many models utilize text alongside audio to maintain accuracy. Moshi uses an "inner monologue" to write out its answer as text immediately before speaking it. Alibaba's Qwen Omni models split the task between a "thinker" that writes the text response and a "talker" that vocalizes it.

    Interaction Modes: Turn-Based vs. Full-Duplex

    • Turn-based: The system waits for the user to finish speaking before generating a response, and stops if the user interrupts.
    • Full-duplex: The model processes two audio streams simultaneously (the user's and its own), allowing it to provide backchannels (like "uh-huh") or handle interruptions naturally.

    Training Methodology

    Using Kyutai's published recipe for Moshi as an example, training a speech-to-speech model involves four main steps:

    1. Training a foundational text model.
    2. Training on millions of hours of raw audio.
    3. Training on conversational data with separated speaker channels.
    4. Fine-tuning on synthetic dialogues generated by a text model and converted to speech.

    Current Limitations

    Despite their capabilities, speech-to-speech models still face several challenges:

    • Reasoning: They generally score lower on complex reasoning tasks compared to text-only models.
    • Tone: Many models struggle to interpret tone correctly when the spoken tone contradicts the literal meaning of the words.
    • Exact Wording: They cannot be fully relied upon to repeat scripts or legal disclaimers word-for-word.
    • Session Limits: Many services impose strict connection or session duration limits.