Videoby Voice AI Space

    Agora's real-time network powers conversational AI and avatars - Voice AI Space Barcelona

    Master real-time networking essentials for powering conversational AI and digital avatars using Agora’s low-latency infrastructure for seamless, interactive user experiences.

    Summary

    About Agora

    Agora is an infrastructure and engagement platform that powers real-time voice, video, and conversational AI experiences. Operating for 12 years, Agora runs a global infrastructure optimized with WebRTC components to deliver sub-second latency worldwide. The company is publicly traded, generates approximately $140 million in annual revenue, and powers 80 billion minutes of audio and video per month.

    The Conversational AI Challenge

    Creating a natural, human-like voice agent requires processing speech, understanding it, and responding within approximately 1.3 seconds. Doing this globally, at scale, on mobile devices, and under challenging network conditions presents significant production hurdles. Agora addresses these challenges by serving as the infrastructure layer for conversational AI.

    The Convo AI Engine and Agent Studio

    Agora has packaged its infrastructure and APIs into the Convo AI Engine, an intelligent orchestration layer that allows enterprises to deploy conversational AI applications with peak performance and high concurrency. For rapid deployment, Agora also offers Agent Studio, a no-code solution. Agora's platform allows developers to mix and match any Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS) technologies without vendor lock-in, optimizing for latency, cost, and unit economics.

    Proprietary Network Infrastructure

    Unlike standard internet routing that relies on least-cost paths, Agora utilizes a proprietary overlay accelerated network. This network dynamically routes data packets over premium paths, enabling ultra-low latency connections (often under 100 milliseconds) between different regions of the world, such as Asia and America.

    Real-Time Avatar Demonstrations

    The platform supports various third-party avatar technologies, which were demonstrated to showcase different use cases:

    • Instant Image-to-Avatar: A model that converts any uploaded image into a speaking avatar within 10 to 12 seconds without prior training, useful for applications like live shopping.
    • Client-Side Avatars: A cost-effective solution that runs entirely within the web browser, eliminating the need for expensive GPU infrastructure.
    • High-Fidelity Lip-Sync Avatars: A highly detailed model suitable for language learning and accessibility, where the visual quality is precise enough to allow lip-reading.

    Agora emphasizes that it does not own the avatar, STT, LLM, or TTS technologies shown; rather, it provides the underlying real-time network and orchestration infrastructure to power them.