Voxtral-Mini-3B-2507
by mistralai
Multimodal audio model for transcription, translation, and voice understanding
mistralai/Voxtral-Mini-3B-2507mixpeek://transcription@v1/mistral_voxtral_mini_3b_v1Overview
Voxtral Mini 3B is Mistral's multimodal audio model combining a Whisper large-v3 encoder with a Ministral-3B language decoder. It handles transcription, translation, audio understanding, and function calling from voice, supporting 8 languages with automatic language detection.
On Mixpeek, Voxtral Mini powers multilingual transcription pipelines and audio understanding workflows. Its ability to answer questions about audio content (not just transcribe) enables richer metadata extraction from podcasts, interviews, and meeting recordings.
Architecture
Three-component architecture: Whisper large-v3 audio encoder (640M) + 4x downsampling audio-language adapter (25M) + Ministral-3B language decoder (3.6B). 32K token context. Handles 30-min audio for transcription, 40-min for understanding. ~9.5 GB GPU RAM in bf16.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so Voxtral-Mini-3B-2507 runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The model produces text, so it lands in payload. Give the
// collection a text vector index and embed that text to make it
// searchable rather than only filterable.
payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
vectors: { "text-embedding": embeddingOfModelOutput },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// universal_extractor@v1 runs google/gemini-embedding-2
// (3072-d) over a bucket, with no inference of your own.Capabilities
- 8-language ASR (EN, ES, FR, PT, HI, DE, NL, IT)
- Automatic language detection
- Audio understanding and question answering
- Function calling from voice input
- Outperforms Whisper large-v3 on Open ASR Leaderboard
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| Open ASR Leaderboard | Mean WER | 7.05% | HuggingFace Open ASR Leaderboard, 2025 |
| LibriSpeech Clean | WER | 1.88% | Mistral, 2025: arxiv,2507.13264 |
| LibriSpeech Other | WER | 4.10% | Mistral, 2025: arxiv,2507.13264 |
Performance
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
Voxtral
arxiv.orgBuild a pipeline with Voxtral-Mini-3B-2507
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free