VibeVoice-ASR-HF
by microsoft
Unified ASR + diarization + timestamps in one 9B model: 60 min single-pass
microsoft/VibeVoice-ASR-HFmixpeek://transcription@v1/microsoft_vibevoice_asr_v1Overview
VibeVoice-ASR is Microsoft's unified speech recognition model that produces structured rich transcriptions (speaker labels, word-level timestamps, and content) from up to 60 minutes of audio in a single forward pass. It replaces the traditional pipeline of separate ASR, diarization, and alignment models with one 9B parameter model.
Supporting 50+ languages with native code-switching (no language flag required), it handles meetings, interviews, podcasts, and call center recordings where knowing who said what matters as much as what was said. On Mixpeek, it powers speaker-attributed transcription for video and audio assets.
Architecture
Encoder-decoder transformer (9B parameters) with multi-task training for simultaneous ASR, speaker diarization, and timestamp alignment. Processes up to 60 minutes of audio in a single pass without sliding window chunking.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so VibeVoice-ASR-HF runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The model produces text, so it lands in payload. Give the
// collection a text vector index and embed that text to make it
// searchable rather than only filterable.
payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
vectors: { "text-embedding": embeddingOfModelOutput },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// universal_extractor@v1 runs google/gemini-embedding-2
// (3072-d) over a bucket, with no inference of your own.Capabilities
- Joint ASR + speaker diarization + word timestamps in one pass
- 60-minute single-pass processing without chunking
- 50+ languages with automatic code-switching
- Structured output: speaker ID, timestamps, and text per segment
- MIT license for unrestricted use
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| Earnings-22 (long-form) | WER | 11.2% | Microsoft, 2026: Model Card |
Performance
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
VibeVoice-ASR: Longform Structured Speech Recognition at Scale
arxiv.orgBuild a pipeline with VibeVoice-ASR-HF
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free