
.wave
Continuous inference engine for running real-time, stateful AI models efficiently.

About .wave
.wave: Continuous inference for real-time AI
.wave is a Y Combinator-backed (W25) platform designed for continuous inference in real-time AI applications. It enables users to run stateful models that process live input throughout a session, offering both a Live API for hosted models and managed deployments for custom models.
Key Features
- .wave Live API: Access hosted models like the NemotronLabs VoiceChat 11B (currently in public beta) and upcoming streaming STT models like Nemotron 3.5 ASR Streaming 0.6B.
- Managed Deployments: Allows users to bring their own continuous models to a managed deployment environment.
- High-Performance Engine: Utilizes WPK, a persistent GPU kernel, to keep recurring work on the GPU, enabling up to 56x more VoiceChat sessions per GPU and lowering GPU cost contributions by 98%.
- Strict Timing Contracts: Ensures ASR streams are kept on schedule with zero missed frames by meeting recurring deadlines for each session.
Use Cases
- Full-Duplex Voice: Running models that require continuous, real-time voice interaction.
- Streaming ASR: Processing live audio streams for automatic speech recognition without falling behind schedule.
Getting Started
Website: https://dotwave.ai
.wave provides a specialized GPU engine built specifically for continuous inference, allowing AI models to keep pace with live, real-time inputs while maximizing GPU capacity and maintaining strict timing schedules.

