Cosmos-Reason2-2B
by nvidia
Physical world reasoning for video understanding
nvidia/Cosmos-Reason2-2Bmixpeek://video_extractor@v1/nvidia_cosmos_reason2_2b_v1Overview
Cosmos-Reason2-2B is NVIDIA's video reasoning model that understands physical interactions, spatial relationships, and causal dynamics in video content. With 2B parameters, it performs temporal reasoning about events, understanding not just what objects are present but how they interact, move, and change over time. The model excels at answering questions about physical plausibility, predicting outcomes, and describing causal chains in video sequences.
Architecture
Vision-language model with temporal attention for video understanding. Uses a visual encoder that processes video frames with temporal position embeddings, feeding into a language model decoder for generating structured descriptions and answering reasoning questions. The architecture includes specialized cross-attention layers for aligning visual frame sequences with language understanding.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so Cosmos-Reason2-2B runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The model produces text, so it lands in payload. Give the
// collection a text vector index and embed that text to make it
// searchable rather than only filterable.
payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
vectors: { "multimodal-embedding": embeddingOfModelOutput },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// universal_extractor@v1 runs google/gemini-embedding-2
// (3072-d) over a bucket, with no inference of your own.Capabilities
- Video temporal reasoning
- Physical interaction understanding
- Spatial relationship analysis
- Causal chain description
- Action prediction and outcome reasoning
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| EgoSchema | Accuracy | 68.4 | Model card |
| PerceptionTest | Accuracy | 62.1 | Model card |
| MVBench | Accuracy | 71.3 | Model card |
Performance
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
Model paper or technical report
arxiv.orgBuild a pipeline with Cosmos-Reason2-2B
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free