NEWVectors or files. Pick a path.Start →
    Models/Captioning/Hcompany/Holo-3.1-4B
    HFScene CaptioningApache 2.0

    Holo-3.1-4B

    by Hcompany

    4B vision-language model for GUI agents and computer-use perception

    1.3Kdl/month
    55likes
    4Bparams
    Identifiers
    Model ID
    Hcompany/Holo-3.1-4B
    Feature URI
    mixpeek://image_extractor@v1/hcompany_holo_31_4b_v1

    Overview

    Holo-3.1-4B is a compact vision-language model tagged for action, agent, computer use, and GUI agents. It is relevant to multimodal search because many agent traces are not documents. They are screenshots, browser states, UI elements, and before-after visual states from tool calls.

    On Mixpeek, Holo can turn screenshots and UI recordings into searchable agent memory. That lets an agent retrieve prior visual states, inspect similar failures, and compare what the screen looked like before deciding whether to retry, stop, or ask for help.

    Architecture

    Qwen-family image-text-to-text model with Hugging Face metadata for action, agent, computer use, GUI agents, and conversational visual reasoning.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so Holo-3.1-4B runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The model produces text, so it lands in payload. Give the
              // collection a text vector index and embed that text to make it
              // searchable rather than only filterable.
              payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
              vectors: { "image-embedding": embeddingOfModelOutput },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // image_extractor@v1 runs google/siglip-base-patch16-224
    // (768-d) over a bucket, with no inference of your own.

    Capabilities

    • GUI and computer-use visual reasoning
    • Screenshot state description for agent memory
    • Compact 4B model size for high-volume UI traces
    • Apache 2.0 licensed model card metadata on Hugging Face

    Use Cases on Mixpeek

    Index browser-agent screenshots and action traces
    Search UI failures by visible state instead of log text
    Compare before-after screens in QA automation
    Give support agents visual memory over prior workflows

    Performance

    Input SizeVariable
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU Memory~10 GB

    Best used with screenshot downsampling and UI event metadata filters.

    Specification

    FrameworkHF
    OrganizationHcompany
    FeatureScene Captioning
    Outputtext
    Modalitiesvideo, image
    RetrieverSemantic Search
    Parameters4B
    LicenseApache 2.0
    Downloads/mo1.3K
    Likes55

    Build a pipeline with Holo-3.1-4B

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free