gemma-4-12B-it
by google
Open 12B multimodal model for image, audio, and long-context agent perception
google/gemma-4-12B-itmixpeek://image_extractor@v1/google_gemma4_12b_it_v1Overview
Gemma 4 12B IT is an instruction-tuned open model from Google DeepMind. The model card describes Gemma 4 as multimodal, with text and image input across the family and audio support on the E2B, E4B, and 12B variants. It is a strong fit for agents that need to inspect retrieved images, short audio clips, or mixed evidence after first-stage search.
On Mixpeek, Gemma 4 12B belongs in the inspection layer. Use cheaper embeddings and filters to retrieve candidates, then ask Gemma to produce concise observations, answer bounded visual questions, or turn multimodal evidence into structured fields that downstream agents can cite.
Architecture
Instruction-tuned Gemma 4 multimodal model exposed through Hugging Face Transformers. The 12B checkpoint supports a 256K context window, multilingual text handling, image input, audio input, and text output.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so gemma-4-12B-it runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The model produces text, so it lands in payload. Give the
// collection a text vector index and embed that text to make it
// searchable rather than only filterable.
payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
vectors: { "multimodal-embedding": embeddingOfModelOutput },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// universal_extractor@v1 runs google/gemini-embedding-2
// (3072-d) over a bucket, with no inference of your own.Capabilities
- Image-text and audio-text understanding in a single instruction-tuned model
- Long-context multimodal reasoning for evidence inspection
- Multilingual support across broad language coverage
- Apache 2.0 license for production evaluation
Use Cases on Mixpeek
Performance
Best used as a second-stage inspector after retrieval narrows candidates
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
Gemma 4 12B IT model card
arxiv.orgBuild a pipeline with gemma-4-12B-it
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free