How Speech-to-Speech AI Models Work: Audio Tokens to Full-Duplex
Master speech to speech AI architecture by exploring audio tokenization and full duplex systems for seamless real time voice interaction.
Summary
Speech-to-Speech Models vs. Three-Model Systems
Traditional voice agents operate like a relay race, using three separate models: speech-to-text to transcribe the input, a language model to generate a text response, and text-to-speech to read the answer aloud. In contrast, a speech-to-speech model processes everything within a single model, taking sound directly as input and producing sound as output without passing a text transcript in between.
How Speech-to-Speech Models Work
Just as text models predict the next token, speech-to-speech models predict the next short slice of sound. To do this, sound must first be converted into discrete pieces using an audio codec. The codec cuts the audio into short slices and describes each slice with numbers from fixed lists called codebooks. For example, Kyutai's Moshi model uses 12.5 slices per second with 8 numbers per slice, where the first number represents what was said and the remaining seven represent how it sounded (voice, tone, and detail).
Hearing, Speaking, and the "Text Trick"
Models process audio input in one of two ways: by reading audio tokens directly or by using an encoder (a listening network) to convert sound into numbers for the language model. Speaking reverses this process, using a decoder to turn predicted audio tokens back into sound. To maintain real-time performance, models like Moshi split the workload between a large model that runs once per slice and a smaller model that fills in the remaining audio numbers.
Because audio tokens spread meaning thinly, many models utilize text alongside audio to maintain accuracy. Moshi uses an "inner monologue" to write out its answer as text immediately before speaking it. Alibaba's Qwen Omni models split the task between a "thinker" that writes the text response and a "talker" that vocalizes it.
Interaction Modes: Turn-Based vs. Full-Duplex
- Turn-based: The system waits for the user to finish speaking before generating a response, and stops if the user interrupts.
- Full-duplex: The model processes two audio streams simultaneously (the user's and its own), allowing it to provide backchannels (like "uh-huh") or handle interruptions naturally.
Training Methodology
Using Kyutai's published recipe for Moshi as an example, training a speech-to-speech model involves four main steps:
- Training a foundational text model.
- Training on millions of hours of raw audio.
- Training on conversational data with separated speaker channels.
- Fine-tuning on synthetic dialogues generated by a text model and converted to speech.
Current Limitations
Despite their capabilities, speech-to-speech models still face several challenges:
- Reasoning: They generally score lower on complex reasoning tasks compared to text-only models.
- Tone: Many models struggle to interpret tone correctly when the spoken tone contradicts the literal meaning of the words.
- Exact Wording: They cannot be fully relied upon to repeat scripts or legal disclaimers word-for-word.
- Session Limits: Many services impose strict connection or session duration limits.
Related Content

On-Device Voice AI Explained: What Runs Locally in 2026

Speaker Diarization Explained: Who Spoke When?

Build or Buy a Voice AI Agent? Frameworks vs Platforms (2026)

Is Your AI Voice Agent Legal? 5 Rules to Know Before You Ship

How Much Does a Voice AI Agent Cost Per Minute? (2026)

How to Test a Voice Agent: Simulated Callers and Evals (2026)

How to Prompt a Voice Agent: Writing for the Ear (2026)

Why Speech-to-Text Gets It Wrong, and How to Fix It (2026)