NEWVectors or files. Pick a path.Start →
    Models/Captioning/nvidia/4D-RGPT-8B
    HFScene CaptioningCC-BY-NC-4.0

    4D-RGPT-8B

    by nvidia

    8B video model for region-grounded 3D and 4D reasoning

    Identifiers
    Model ID
    nvidia/4D-RGPT-8B
    Feature URI
    mixpeek://video_extractor@v1/nvidia_4d_rgpt_8b_v1

    Overview

    4D-RGPT-8B is an NVIDIA video-text model focused on region grounding, 3D reasoning, and 4D reasoning. Those capabilities are important when an agent needs more than a clip-level summary. The agent needs to know which region changed, where the object moved, and how the event evolved over time.

    On Mixpeek, 4D-RGPT can enrich video indexes with region-grounded temporal evidence. It is a fit for robotics footage, surveillance review, sports clips, and operational video where the retrieval result must preserve spatial and temporal context.

    Architecture

    NVILA-Lite-8B based video-text-to-text model. The Hugging Face metadata tags it for video understanding, region grounding, 3D reasoning, 4D reasoning, and perceptual distillation.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so 4D-RGPT-8B runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The model produces text, so it lands in payload. Give the
              // collection a text vector index and embed that text to make it
              // searchable rather than only filterable.
              payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
              vectors: { "multimodal-embedding": embeddingOfModelOutput },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // universal_extractor@v1 runs google/gemini-embedding-2
    // (3072-d) over a bucket, with no inference of your own.

    Capabilities

    • Region-grounded video understanding
    • 3D and 4D reasoning over spatial-temporal evidence
    • Video-text-to-text analysis for agent perception loops
    • Designed for grounding objects and events through time

    Use Cases on Mixpeek

    Retrieve clips where an object moves through a specific region
    Audit robotics or physical-world agent observations
    Build evidence bundles for security and operations review
    Search sports or live-event video with spatial-temporal constraints

    Performance

    Input SizeVariable
    GPU LatencyInput dependent
    GPU ThroughputVideo length dependent
    GPU Memory~18 GB

    Region-grounded video reasoning cost depends heavily on clip length and frame sampling.

    Specification

    FrameworkHF
    Organizationnvidia
    FeatureScene Captioning
    Outputtext
    Modalitiesvideo, image
    RetrieverSemantic Search
    Parameters8B
    LicenseCC-BY-NC-4.0
    Downloads/mo108
    Likes13

    Research Paper

    4D-RGPT

    arxiv.org

    Build a pipeline with 4D-RGPT-8B

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free