NEWVectors or files. Pick a path.Start →
    Models/Speech & Audio/usefulsensors/moonshine-streaming-medium
    HFTranscriptionMIT

    moonshine-streaming-medium

    by usefulsensors

    245M streaming ASR with 107ms latency: beats Whisper Large V3 at 6x fewer parameters

    Identifiers
    Model ID
    usefulsensors/moonshine-streaming-medium
    Feature URI
    mixpeek://transcription@v1/moonshine_streaming_medium_v1

    Overview

    Moonshine Streaming Medium is a 245M-parameter automatic speech recognition model designed for real-time, low-latency streaming on edge-class hardware. It pairs a lightweight 50Hz audio frontend with a sliding-window Transformer encoder that uses bounded local attention and no positional embeddings (an "ergodic" encoder), while an adapter injects positional information before a standard autoregressive decoder.

    Trained on roughly 300K hours of speech data, the model achieves transcription quality on par with Whisper Large V3 while running at 107ms latency on a MacBook Pro and using 6x fewer parameters. On Mixpeek, Moonshine Streaming provides a fast, lightweight alternative to Whisper for English ASR pipelines where latency and compute cost matter more than multilingual support.

    Architecture

    Lightweight 50Hz audio frontend + sliding-window Transformer encoder with bounded local attention and no positional embeddings (ergodic encoder). Adapter layer injects positional information before autoregressive decoder. 245M total parameters. Trained on ~300K hours of speech data.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so moonshine-streaming-medium runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The model produces text, so it lands in payload. Give the
              // collection a text vector index and embed that text to make it
              // searchable rather than only filterable.
              payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
              vectors: { "text-embedding": embeddingOfModelOutput },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // universal_extractor@v1 runs google/gemini-embedding-2
    // (3072-d) over a bucket, with no inference of your own.

    Capabilities

    • 107ms streaming latency on consumer hardware
    • Accuracy matching Whisper Large V3 at 6x fewer params
    • Ergodic encoder for unbounded-length streaming
    • Optimized for edge and on-device deployment
    • 245M parameters: fits on mobile and embedded hardware

    Use Cases on Mixpeek

    Real-time video transcription: stream captions during live content ingestion
    Edge ASR pipelines: transcribe audio on-device before uploading to Mixpeek
    Low-latency content indexing: process audio streams with minimal delay for near-real-time search

    Benchmarks

    DatasetMetricScoreSource
    LibriSpeech (clean)WER~3.0%Useful Sensors, 2026: arxiv,2602.12241
    Edge latency (MacBook Pro)Latency107msUseful Sensors, 2026: arxiv,2602.12241
    vs Whisper Large V3Params ratio6x smaller, comparable WERUseful Sensors, 2026: arxiv,2602.12241

    Performance

    Input SizeStreaming audio (unbounded length)
    GPU Latency~50ms / chunk (A100)
    CPU Latency~107ms / chunk (MacBook Pro)
    GPU Throughput~20x real-time (A100)
    GPU Memory~0.8 GB

    Specification

    FrameworkHF
    Organizationusefulsensors
    FeatureTranscription
    Outputtext + timestamps
    Modalitiesvideo, audio
    RetrieverTranscript Search
    Parameters245M
    LicenseMIT
    Downloads/mo180K

    Research Paper

    Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications

    arxiv.org

    Build a pipeline with moonshine-streaming-medium

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free