NEWVectors or files. Pick a path.Start →
    Models/Captioning/google/gemma-4-31B-it
    HFScene CaptioningApache-2.0

    gemma-4-31B-it

    by google

    Top-3 open VLM with 256K context for dense visual document understanding

    Identifiers
    Model ID
    google/gemma-4-31B-it
    Feature URI
    mixpeek://image_extractor@v1/google_gemma4_31b_v1

    Overview

    Gemma 4 31B is Google's dense vision-language model, currently ranked #3 among open models on the Arena AI text leaderboard. Unlike the MoE variant (27B-A4B), this dense model activates all 31B parameters, delivering the highest quality at higher compute cost.

    The 256K context window and built-in thinking mode make it particularly strong for complex document understanding tasks where accuracy matters more than throughput.

    Architecture

    Dense transformer architecture with 31B parameters. Vision encoder processes image patches. 256K context window. Thinking mode enables chain-of-thought reasoning for complex visual tasks.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so gemma-4-31B-it runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The model produces text, so it lands in payload. Give the
              // collection a text vector index and embed that text to make it
              // searchable rather than only filterable.
              payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
              vectors: { "multimodal-embedding": embeddingOfModelOutput },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // universal_extractor@v1 runs google/gemini-embedding-2
    // (3072-d) over a bucket, with no inference of your own.

    Capabilities

    • Highest-quality open VLM (Arena #3)
    • 256K context window
    • Dense architecture for fine-tuning
    • Built-in reasoning mode
    • Apache 2.0 license

    Use Cases on Mixpeek

    High-accuracy visual document extraction where quality is critical
    Complex chart and diagram understanding
    Fine-tuning on domain-specific visual data (dense architecture)

    Benchmarks

    DatasetMetricScoreSource
    MMLU ProAccuracy85.2%Google, May 2026
    AIME 2026Accuracy89.2%Google, May 2026
    Arena AI LeaderboardELOTop 3 openArena AI, May 2026

    Performance

    Input SizeUp to 256K tokens (text + image patches)
    GPU Latency~280ms / image (A100)
    GPU Throughput~28 images/sec (A100, batch 4)
    GPU Memory~62 GB (dense, full activation)

    Specification

    FrameworkHF
    Organizationgoogle
    FeatureScene Captioning
    Outputtext
    Modalitiesvideo, image
    RetrieverSemantic Search
    Parameters31B
    LicenseApache-2.0
    Downloads/mo820K

    Research Paper

    Gemma 4: Byte for byte, the most capable open models

    arxiv.org

    Build a pipeline with gemma-4-31B-it

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free