NEWVectors or files. Pick a path.Start →
    Models/Speech & Audio/Qwen/Qwen3-ForcedAligner-0.6B
    HFTranscriptionApache 2.0

    Qwen3-ForcedAligner-0.6B

    by Qwen

    Forced alignment model for word and character timestamps in agent audio search

    Identifiers
    Model ID
    Qwen/Qwen3-ForcedAligner-0.6B
    Feature URI
    mixpeek://transcription@v1/qwen3_forced_aligner_06b_v1

    Overview

    Qwen3-ForcedAligner-0.6B aligns known text to speech and predicts timestamps for words or characters. It is part of the Qwen3-ASR release and is designed to add precise timing to transcripts, which is critical when an agent has to cite the exact spoken evidence behind an answer.

    On Mixpeek, forced alignment turns raw transcript text into searchable spans with start and end times. Agents can retrieve a sentence, jump to the correct audio or video moment, and cite the matching time range instead of returning an ungrounded transcript blob.

    Architecture

    LLM-based non-autoregressive timestamp predictor from the Qwen3-ASR family. The model aligns text-speech pairs and supports timestamp prediction for arbitrary units within speech windows.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so Qwen3-ForcedAligner-0.6B runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The model produces text, so it lands in payload. Give the
              // collection a text vector index and embed that text to make it
              // searchable rather than only filterable.
              payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
              vectors: { "text-embedding": embeddingOfModelOutput },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // universal_extractor@v1 runs google/gemini-embedding-2
    // (3072-d) over a bucket, with no inference of your own.

    Capabilities

    • Word-level and character-level timestamp prediction
    • Alignment for Chinese, English, Cantonese, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish
    • Pairs with Qwen3-ASR transcription models
    • Useful for citation-grade audio and video retrieval

    Use Cases on Mixpeek

    Attach exact timestamps to call center and meeting transcripts
    Build video agents that answer with cited spoken evidence
    Index subtitle-quality transcript spans for semantic search
    Repair coarse ASR timestamps before chunking and retrieval

    Benchmarks

    DatasetMetricScoreSource
    Qwen3-ASR timestamp evalsTimestamp accuracy-Qwen3-ASR model card

    Performance

    Input SizeUp to 5 minutes of speech per alignment window
    GPU LatencyAudio duration dependent
    GPU ThroughputBatch dependent
    GPU Memory~2 GB

    Use after transcription when retrieval needs word-level citation

    Specification

    FrameworkHF
    OrganizationQwen
    FeatureTranscription
    Outputtext + timestamps
    Modalitiesvideo, audio
    RetrieverTranscript Search
    Parameters0.6B
    LicenseApache 2.0
    Downloads/mo404K

    Research Paper

    Qwen3-ASR Technical Report

    arxiv.org

    Build a pipeline with Qwen3-ForcedAligner-0.6B

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free