SenseNova-U1-8B-MoT
by sensenova
8B any-to-any multimodal model for image understanding, generation, and editing
sensenova/SenseNova-U1-8B-MoTmixpeek://image_extractor@v1/sensenova_u1_8b_mot_v1Overview
SenseNova-U1-8B-MoT is an any-to-any multimodal model tagged for feature extraction, image-to-text, text-to-image, image editing, and custom-code inference. That mix matters for agents because perception is often not a single captioning call: an agent may need to inspect an image, generate an explanation, propose an edit, and preserve evidence of what changed.
On Mixpeek, SenseNova U1 fits pipelines that retrieve visual evidence first, then ask a multimodal model to explain or transform that evidence. It is especially relevant for creative QA, ad review, product imagery, and human-in-the-loop visual analysis.
Architecture
8B-class mixture-of-transformers style any-to-any multimodal model. Supports image-to-text, text-to-image, image editing, and feature extraction paths according to the Hugging Face model metadata.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so SenseNova-U1-8B-MoT runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The model produces text, so it lands in payload. Give the
// collection a text vector index and embed that text to make it
// searchable rather than only filterable.
payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
vectors: { "image-embedding": embeddingOfModelOutput },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// image_extractor@v1 runs google/siglip-base-patch16-224
// (768-d) over a bucket, with no inference of your own.Capabilities
- Any-to-any multimodal interaction across image and text tasks
- Image-to-text reasoning for visual evidence review
- Text-to-image and image-editing paths for iterative agent workflows
- Apache 2.0 licensed model card metadata on Hugging Face
Use Cases on Mixpeek
Performance
Any-to-any models should be routed to the narrowest task path needed for the agent step.
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
SenseNova-U1
arxiv.orgBuild a pipeline with SenseNova-U1-8B-MoT
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free