NEWVectors or files. Pick a path.Start →
    Models/Speech & Audio/pyannote/speaker-diarization-community-1
    HFSpeaker Diarizationcc-by-4.0

    speaker-diarization-community-1

    by pyannote

    Community speaker diarization pipeline for who-spoke-when audio metadata

    5.1Mdl/month
    1,156likes
    Pipelineparams
    Identifiers
    Model ID
    pyannote/speaker-diarization-community-1
    Feature URI
    mixpeek://transcription@v1/pyannote_diarization_community_1

    Overview

    pyannote Community-1 is a speaker diarization pipeline that segments audio by speaker turns, speech activity, speaker changes, and overlapped speech. It is publicly accessible with license acceptance and has become one of the highest-traffic diarization models on HuggingFace.

    On Mixpeek, diarization turns raw audio and video transcripts into searchable conversational structure. Agents can ask not only what was said, but who said it and when it happened.

    Architecture

    pyannote.audio pipeline composed of voice activity detection, speaker change detection, overlapped speech detection, embedding, and clustering components. Accepts whole files or waveform excerpts.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so speaker-diarization-community-1 runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The model produces text, so it lands in payload. Give the
              // collection a text vector index and embed that text to make it
              // searchable rather than only filterable.
              payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
              vectors: { "text-embedding": embeddingOfModelOutput },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // universal_extractor@v1 runs google/gemini-embedding-2
    // (3072-d) over a bucket, with no inference of your own.

    Capabilities

    • Speaker turn segmentation
    • Voice activity and speaker change detection
    • Overlapped speech handling
    • Runs through pyannote.audio

    Use Cases on Mixpeek

    Search meeting recordings by speaker and topic
    Build agent memory over multi-speaker calls
    Filter podcast or interview clips by host, guest, or caller
    Attach who-spoke-when metadata to transcripts for audit trails

    Performance

    Input SizeAudio file or waveform excerpt
    GPU LatencyAudio duration dependent
    GPU ThroughputBatch by file for offline archives
    GPU Memory~2 GB

    Model files require accepting HuggingFace access conditions

    Specification

    FrameworkHF
    Organizationpyannote
    FeatureSpeaker Diarization
    Outputspeaker segments
    Modalitiesvideo, audio
    RetrieverSpeaker Filter
    ParametersPipeline
    Licensecc-by-4.0
    Downloads/mo5.1M
    Likes1,156

    Research Paper

    pyannote.audio speaker diarization community-1

    arxiv.org

    Build a pipeline with speaker-diarization-community-1

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free