NEWVectors or files. Pick a path.Start →
    Models/Captioning/nvidia/Cosmos-Reason2-2B
    HFScene CaptioningNVIDIA Open Model License

    Cosmos-Reason2-2B

    by nvidia

    Physical world reasoning for video understanding

    Identifiers
    Model ID
    nvidia/Cosmos-Reason2-2B
    Feature URI
    mixpeek://video_extractor@v1/nvidia_cosmos_reason2_2b_v1

    Overview

    Cosmos-Reason2-2B is NVIDIA's video reasoning model that understands physical interactions, spatial relationships, and causal dynamics in video content. With 2B parameters, it performs temporal reasoning about events, understanding not just what objects are present but how they interact, move, and change over time. The model excels at answering questions about physical plausibility, predicting outcomes, and describing causal chains in video sequences.

    Architecture

    Vision-language model with temporal attention for video understanding. Uses a visual encoder that processes video frames with temporal position embeddings, feeding into a language model decoder for generating structured descriptions and answering reasoning questions. The architecture includes specialized cross-attention layers for aligning visual frame sequences with language understanding.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so Cosmos-Reason2-2B runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The model produces text, so it lands in payload. Give the
              // collection a text vector index and embed that text to make it
              // searchable rather than only filterable.
              payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
              vectors: { "multimodal-embedding": embeddingOfModelOutput },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // universal_extractor@v1 runs google/gemini-embedding-2
    // (3072-d) over a bucket, with no inference of your own.

    Capabilities

    • Video temporal reasoning
    • Physical interaction understanding
    • Spatial relationship analysis
    • Causal chain description
    • Action prediction and outcome reasoning

    Use Cases on Mixpeek

    Manufacturing quality control, detecting anomalous physical interactions
    Autonomous driving scene understanding
    Sports analytics with play-by-play reasoning
    Surveillance video summarization with causal explanations

    Benchmarks

    DatasetMetricScoreSource
    EgoSchemaAccuracy68.4Model card
    PerceptionTestAccuracy62.1Model card
    MVBenchAccuracy71.3Model card

    Performance

    Input SizeVariable
    GPU Latency~250ms per 16-frame clip on A100
    GPU Throughput~4 clips/sec
    GPU MemoryModel dependent

    Specification

    FrameworkHF
    Organizationnvidia
    FeatureScene Captioning
    Outputtext
    Modalitiesvideo, image
    RetrieverSemantic Search
    Parameters2B
    LicenseNVIDIA Open Model License
    Downloads/mo169K

    Research Paper

    Model paper or technical report

    arxiv.org

    Build a pipeline with Cosmos-Reason2-2B

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free