speaker-diarization-community-1
by pyannote
Community speaker diarization pipeline for who-spoke-when audio metadata
pyannote/speaker-diarization-community-1mixpeek://transcription@v1/pyannote_diarization_community_1Overview
pyannote Community-1 is a speaker diarization pipeline that segments audio by speaker turns, speech activity, speaker changes, and overlapped speech. It is publicly accessible with license acceptance and has become one of the highest-traffic diarization models on HuggingFace.
On Mixpeek, diarization turns raw audio and video transcripts into searchable conversational structure. Agents can ask not only what was said, but who said it and when it happened.
Architecture
pyannote.audio pipeline composed of voice activity detection, speaker change detection, overlapped speech detection, embedding, and clustering components. Accepts whole files or waveform excerpts.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so speaker-diarization-community-1 runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The model produces text, so it lands in payload. Give the
// collection a text vector index and embed that text to make it
// searchable rather than only filterable.
payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
vectors: { "text-embedding": embeddingOfModelOutput },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// universal_extractor@v1 runs google/gemini-embedding-2
// (3072-d) over a bucket, with no inference of your own.Capabilities
- Speaker turn segmentation
- Voice activity and speaker change detection
- Overlapped speech handling
- Runs through pyannote.audio
Use Cases on Mixpeek
Performance
Model files require accepting HuggingFace access conditions
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
pyannote.audio speaker diarization community-1
arxiv.orgBuild a pipeline with speaker-diarization-community-1
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free