Qwen3-ForcedAligner-0.6B
by Qwen
Forced alignment model for word and character timestamps in agent audio search
Qwen/Qwen3-ForcedAligner-0.6Bmixpeek://transcription@v1/qwen3_forced_aligner_06b_v1Overview
Qwen3-ForcedAligner-0.6B aligns known text to speech and predicts timestamps for words or characters. It is part of the Qwen3-ASR release and is designed to add precise timing to transcripts, which is critical when an agent has to cite the exact spoken evidence behind an answer.
On Mixpeek, forced alignment turns raw transcript text into searchable spans with start and end times. Agents can retrieve a sentence, jump to the correct audio or video moment, and cite the matching time range instead of returning an ungrounded transcript blob.
Architecture
LLM-based non-autoregressive timestamp predictor from the Qwen3-ASR family. The model aligns text-speech pairs and supports timestamp prediction for arbitrary units within speech windows.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so Qwen3-ForcedAligner-0.6B runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The model produces text, so it lands in payload. Give the
// collection a text vector index and embed that text to make it
// searchable rather than only filterable.
payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
vectors: { "text-embedding": embeddingOfModelOutput },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// universal_extractor@v1 runs google/gemini-embedding-2
// (3072-d) over a bucket, with no inference of your own.Capabilities
- Word-level and character-level timestamp prediction
- Alignment for Chinese, English, Cantonese, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish
- Pairs with Qwen3-ASR transcription models
- Useful for citation-grade audio and video retrieval
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| Qwen3-ASR timestamp evals | Timestamp accuracy | - | Qwen3-ASR model card |
Performance
Use after transcription when retrieval needs word-level citation
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
Qwen3-ASR Technical Report
arxiv.orgBuild a pipeline with Qwen3-ForcedAligner-0.6B
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free