Soniox is launching TTS v2, which it calls its most powerful text-to-speech model to date. Its available in the United States, Europe, and Japan. One model handles voice quality, expressive control, voice cloning, precise pronunciation, and generation in more than 60 languages, including switching between them mid-utterance, all with ultra-low-latency streaming.
Soniox aims it at voice agents, global consumer products, interactive characters, accessibility tools, games, media applications, and other real-time voice experiences. The company built the whole system in-house including model architecture, audio codec, and inference engine.
Try it out yourself
One model instead of several
The pitch is that changing language or use case does not require a different model or a different voice. According to Soniox, the same voice delivers expressive speech, follows directions about emotion and delivery, speaks more than 60 languages, and switches between them within a sentence without losing its identity. It also pronounces complex information accurately, streams in real time, stops cleanly when interrupted, and resumes from the correct point. All of that is meant to hold at production scale and cost.
Expressive control through audio tags
Audio tags let developers direct how a voice performs the text, controlling emotion, tone, delivery, and vocal reactions. Tags can be inserted throughout a passage instead of fixing one style for the whole piece:
[whispering] Shh, shh — she's coming. [pause] [excited] SURPRISE!!! [loudly] Happy birthday!!! [laughing] Look at her face!
With no tags, the model adapts its rhythm, pacing, emphasis, pauses, and expression to the meaning of the text.
60+ languages, with mid-sentence switching
TTS v2 was multilingual from the start. The voice model keeps its identity, character, pronunciation quality, and expressive range as it moves between languages, and that it can follow a language change inside a single continuous utterance. That covers foreign names, international addresses, product names, or a full phrase in another language:
Your reservation is confirmed. Cuando llegues al hotel, muestra este código en recepción: ES-4928.
Voice cloning from seconds of audio
TTS v2 produces what Soniox describes as a high-fidelity voice clone from seconds of audio, capturing voice identity, accent, rhythm, pacing, personality, delivery, and expressive range. Everyday recordings work too, Soniox strips background noise, echo, and recording artifacts from the source, including samples recorded on a phone or outside a studio. Voice cloning must only be used with the rights and consent required to create and use the voice.
Precision for names, numbers, and codes
The model is meant to preserve the text and speak complex information precisely. That includes people and place names; medical, scientific, legal, financial, and technical terminology; phone numbers, email addresses, prices, dates, and addresses; and verification codes, account numbers, and other alphanumeric identifiers.
Your verification code is 7Q4M9B. I repeat: 7Q4M9B. Your case number is CX-8047-A19, and the affected device is model XR-12 Pro. Your total is $1,284.37, including tax and delivery.
Built for real-time voice agents
Ultra-low-latency streaming starts audio playback while the input text is still being generated, so the model does not wait for a complete sentence or response. That matters when text arrives progressively from a language model. Character-level timestamps let applications highlight text as it is spoken, track how much a user has heard, stop playback cleanly on interruption, avoid repeating spoken content, resume from the correct point, and sync speech with captions or animation. Reduced silence between sentences and around punctuation keeps conversations responsive.
Pricing
TTS v2 costs $0.70 per generated hour. That rate includes every feature: expressive generation, more than 60 languages, audio-tag control, built-in and cloned voices, language switching, precision for structured information, ultra-low-latency streaming, and character-level timestamps.
Availability and upgrading
The model, voices, API, and capabilities are the same in the United States, Europe, and Japan, with regional processing and low-latency streaming in every region. The API is backward compatible with the previous version, so existing users upgrade by changing the model name from tts-rt-v1 to tts-rt-v2.
About Soniox
Soniox builds foundational voice AI for real-time human communication.
The company develops its models and voice infrastructure from scratch, with a focus on multilingual speech, accuracy, low latency, production reliability, and global deployment.
Soniox technology is used by companies building voice agents, transcription products, healthcare applications, communication tools, accessibility products, wearables, and other voice-enabled experiences.
Soniox powers many of the well known brands including Perplexity, Samsung, LG, HappyRobot, and Wonderful. You can find more about the customers on their wesbite soniox.com.
Important links:
Platform: https://soniox.com/text-to-speech
Benchmarks mentioning Soniox:
Speech-to-Text: https://github.com/pipecat-ai/stt-benchmark (Pipecat, Daily.co)
Text-to-Speech: https://benchmarks.coval.ai/tts (Coval)
Product launch · Published in partnership with Soniox



