Find all the Voice AI Startup programs:
    Videoby Voice AI Space

    Speaker Diarization Explained: Who Spoke When?

    Master speaker diarization to identify unique voices in audio, segmenting speech by individual to enhance automated transcription and conversation analysis.

    Summary

    What is Speaker Diarization?

    Speaker diarization is the process of determining "who spoke when" in an audio recording, assigning anonymous labels (such as Speaker 0, Speaker 1, etc.) to different speakers. It is distinct from speaker identification, which matches a voice to a specific known person using a reference sample. Diarization produces a timeline of speaker turns that can be aligned with a text transcript.

    How Diarization Works

    There are two primary approaches to speaker diarization:

    • The Classic Pipeline: This method involves three steps:
      1. Voice Activity Detection (VAD): Finding and isolating speech while removing silence and noise, then cutting the speech into short segments.
      2. Speaker Embedding: A model analyzes each segment to generate a "voiceprint" (a list of numbers representing the characteristics of the voice).
      3. Clustering: Grouping similar voiceprints together. The system must estimate the number of speakers; grouping too loosely merges different speakers, while grouping too strictly splits a single speaker's turns.
    • End-to-End Models: A single neural network processes the audio and directly outputs who is speaking frame-by-frame. This approach natively handles overlapping speech (when multiple people talk at once) but is typically limited to a fixed maximum number of speaker slots (such as 4 or 8).

    Open-Source Examples

    • NVIDIA NeMoTron 3 Diarization: An end-to-end model supporting up to 8 speakers. It can run offline or streaming (with a recommended delay of about a third of a second) and uses open weights.
    • PyAnnote (Community-1): A pipeline-based model that uses segmentation, voiceprints, and clustering. It employs a small end-to-end model in its first step to detect overlapping speech. It is designed for complete recordings.

    Measuring Accuracy and Common Challenges

    Diarization accuracy is measured using the Diarization Error Rate (DER), which calculates the total time of missed speech, false alarms, and speaker confusion divided by the total reference speech time. Diarization systems commonly struggle with:

    • Overlapping speech
    • Quick speaker turns (within 0.5 seconds of a change)
    • Similar-sounding voices
    • Far-away or table microphones compared to close-up headsets
    • Recordings with a high number of speakers

    Relevance for Voice Agents

    Voice agents often do not require diarization if the agent and the caller are recorded on separate audio channels. Diarization becomes necessary when multiple speakers share a single channel (such as on a speakerphone or in a meeting). When implementing diarization, developers should note that streaming diarization is generally less accurate than processing a full recording, and voiceprints may be classified as biometric data requiring explicit user consent under regulations like GDPR.