Find any scene in your video library.
Mixpeek understands your video, image, audio, and documents, and returns the most accurate, timestamped results from the storage you already have. Every search teaches it: relevance keeps improving the more you and your agents use it.



Explore the live demos above with zero setup: then start here when you want Mixpeek running on your own data.
Bring your own vectors
You already have embeddingsMixpeek Vector Store (MVS): an agent-native vector store that runs on your object storage. Dense, sparse, and BM25 search. From $25/mo.
Connect your files
You have raw filesManaged indexing: point us at video, audio, or documents and we extract scenes, faces, OCR, transcripts, and embeddings. No pipeline to build. From $25/mo.
What the agent sees.
Every object you index becomes structured, searchable features: faces and objects in a frame, layout regions in a document, speakers in audio. Those are the same features an agent queries, and joins across modalities.
Hover or tap a card to preview the search it powers.
Person · 0.97Face · 0.95Handbag · 0.92Video · 00:04:12“A woman carrying a tan tote bag walks past a red storefront.”
transcript · “…meet me at the corner in five.”
HeaderChartBodySignaturePDF · resume.pdfHeader, body, charts, and signature detected as typed regions.
OCR · 1 header · 3 sections · 1 signature

Who spoke when, aligned to the transcript and the timeline.
audio · 2 speakers · matched at 00:01:30
One query across every modality.
Real questions rarely fit one feature. “Find the moment our CEO said guidance while the slide read Q4 outlook” needs a face, a spoken phrase, and on-screen text to line up at the same instant.
Mixpeek ties those features to the same object and timestamp, so an agent gets back the exact clip instead of three unrelated matches.
The CEO says “guidance” as the slide behind her reads “Q4 outlook.”
"What concepts exist in my data that nobody has labeled yet?"
No keyword, example, or prompt can answer that: they all assume you already know what you're looking for. Mixpeek clusters what belongs together on its own, so a natural hierarchy surfaces instead of a flat pile of tags: a taxonomy built from your data, organized around what your business actually cares about.
That taxonomy is your ground truth, and it feeds back. Every search, every correction sharpens the features, the clusters, and the relationships. Your competitors' metadata decays. Yours compounds.
- Creative moments644
- └Unboxing214
- └Hands-on close-up121
- └Reveal + reaction93
- └Night driving88
- └Product on white342
every search + correction feeds back → sharper features, tighter clusters, truer taxonomy
"Have I solved this before?"
A computer-use agent hits a screen it cannot get past. Somebody unsticks it once. Thirty days later the same wall appears on a redesigned page, under a different account, and the agent recovers on its own by retrieving what worked last time. Not a summary of the conversation: the screen, the action, the tool result, and the transition that followed.
Retrieval returns the packet itself, so the agent can read the evidence rather than trust a paraphrase of it.
{
"observation_id": "obs_run_42_step_08",
"run_id": "run_42",
"step": 8,
"task": "reset account password",
"action_before": { "type": "click", "target": "Save" },
"visual_state": {
"caption": "Settings page, save button disabled",
"ocr": ["Current password", "New password", "Save"],
"ui_elements": [
{ "role": "button", "text": "Save", "enabled": false }
]
},
"tool_outputs": [
{ "tool": "validate_password_policy",
"result": "missing_special_character" }
]
}- Screen · a settings page with the save control greyed out
- OCR · "Current password", "New password", "Save"
- UI state · button "Save", enabled: false
- Tool output · validate_password_policy → missing_special_character
- step_09added a special character to the new password
- step_10Save became enabled
- step_11account updated
Ask the agent why it did that and the answer is a citation, not a recollection: the run, the step, the screen it matched, and the transition it copied.
Start from a complete namespace.
Use a template when the outcome is clear and the stack should arrive together: storage sources, extraction stages, retrievers, evaluation rules, and deployment notes.
Sources, buckets, collections, and retrievers ship together.
One YAML file mirrors the rendered flow diagram.
Thresholds and review rules are attached to the template.
One install. Two paths.
Most retrieval stacks mean gluing together a vector DB, a file pipeline, and an agent layer. Mixpeek is one install with two ways in.
Bring embeddings
Plugs into your existing stack.
Connect your storage, point Mixpeek at it, and every file becomes searchable by what's inside it. No migration, no code changes.

Mux
Every Mux upload becomes searchable by face, scene, transcript, and on-screen text, with no manual tagging.
View integration →
Backblaze B2
S3-compatible extraction at 1/5th the cost. Store on B2, extract with Mixpeek, zero egress fees.
View integration →Iconik
Every asset in your DAM becomes findable by what's inside it: scenes, faces, spoken words, on-screen text.
View integration →Pick your file types. Choose what to search by.
Video, image, audio, documents, text, or web: connect a bucket, pick the features you want to search by, and these pipelines run as they are. Every one is documented and open source in the extractor cookbook.
Search video by
7 live extractors
- Scenes, speech & visual similarityMultimodal (Video/Audio/Image) Try it live Docs
- Everything in one pass, any fileUniversal All-in-One Try it live Docs
- Whole objects: all their files as oneMulti-File Object Embeddings (Gemini) Try it live Docs
- The same face, across your whole libraryFace Identity (SCRFD + ArcFace) Try it live Docs
- Sounds and music, by how they soundAudio Fingerprinting (CLAP) Try it live Docs
- Text that moves across the screenScrolling/Marquee Text OCR Try it live Docs
- Your metadata, with zero computePassthrough (Storage Only) Try it live Docs
From $25/mo. Usage-based everything.
Two products, one model: a monthly minimum that acts as a floor, with usage above it billed at the same transparent rates. MVS is priced by the vector, Managed by the object.
Bring your own embeddings and pay by the vector. Dense, sparse, and BM25 search on your own object storage. Build starts at $25/mo with up to 1M vectors; Scale ($250/mo) covers 25M.
Start with MVSBring raw objects and pay by the object: credits at $0.001 cover extraction, embedding, indexing, enrichment, and retrieval. Build covers 100K objects/mo; Scale ($250/mo) covers 1M.
Start with ManagedDedicated infrastructure, self-hosted options, SSO, SLA, security reviews, and hands-on architecture support.
Talk to usCommon questions.
Do I have to move my data?
No. Mixpeek reads from your existing S3, GCS, R2, Azure, or S3-compatible bucket. Your storage stays the system of record, and nothing leaves your cloud.
How fast is retrieval?
Hybrid queries (dense, sparse, and BM25) return in well under 100ms p95, even with vectors persisted on object storage rather than held in RAM.
Do I need embeddings to start?
No. Bring your own vectors with MVS, or point Managed at raw files and it generates embeddings and features for you.
What can Managed extract?
Faces, scenes, transcripts, OCR, labels, and embeddings from video, images, audio, PDFs, and documents, all indexed at the object level.
Can I self-host?
Yes. Deploy in your own cloud (BYO-Cloud) with encryption, role-based access, SSO, audit trails, and namespaces. We hold no SOC 2 or HIPAA certification today; see mixpeek.com/trust.
How does pricing work?
Both MVS and Managed start at $25/mo minimum. Usage counts toward the minimum: pay the greater of metered usage or the floor. MVS bills storage + queries; Managed bills in credits covering extraction, embedding, indexing, and retriever execution.








