NEWVectors or files. Pick a path.Start →
    Models/Embeddings/mixedbread-ai/mxbai-colbert-large-v1
    HFText EmbeddingsApache 2.0

    mxbai-colbert-large-v1

    by mixedbread-ai

    Late-interaction ColBERT model for high-recall token-level retrieval

    Identifiers
    Model ID
    mixedbread-ai/mxbai-colbert-large-v1
    Feature URI
    mixpeek://text_extractor@v1/mxbai_colbert_large_v1

    Overview

    mxbai-colbert-large-v1 from mixedbread.ai is a ColBERT-style late-interaction retrieval model that produces per-token embeddings for both queries and documents. Instead of compressing an entire passage into a single vector, it retains token-level representations and computes relevance via MaxSim: the maximum similarity between each query token and all document tokens. This architecture captures fine-grained lexical and semantic matches that single-vector models miss.

    Architecture

    Late-interaction transformer based on the ColBERT architecture. Encodes queries and documents independently into per-token embeddings, then scores relevance using MaxSim: for each query token, find the maximum cosine similarity with any document token, then sum across query tokens. This allows pre-computation of document embeddings while retaining token-level matching at query time.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so mxbai-colbert-large-v1 runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The vector name has to match a vector index on the collection.
              vectors: { "text-embedding": yourVector },
              payload: { source_key: "archive/2026/asset-00412" },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // text_extractor@v1 runs intfloat/multilingual-e5-large-instruct
    // (1024-d) over a bucket, with no inference of your own.

    Capabilities

    • Token-level semantic matching
    • High-recall retrieval
    • Fine-grained relevance scoring
    • Efficient pre-computed document indexing

    Use Cases on Mixpeek

    First-stage retrieval where recall matters more than latency
    Legal discovery with precise term matching requirements
    Scientific literature search across technical terminology
    Question answering over large document collections

    Benchmarks

    DatasetMetricScoreSource
    BEIR (avg)nDCG@1056.2Model card
    MS MARCO DevMRR@1040.8Model card
    LoTTESuccess@578.3Model card

    Performance

    Input SizeVariable
    GPU Latency~12ms per query on A100 (pre-indexed)
    GPU Throughput~80 queries/sec
    GPU MemoryModel dependent

    Specification

    FrameworkHF
    Organizationmixedbread-ai
    FeatureText Embeddings
    Output1024-dim vector
    Modalitiesdocument, audio
    RetrieverText Similarity
    Parameters335M
    LicenseApache 2.0
    Downloads/mo22K

    Research Paper

    Model paper or technical report

    arxiv.org

    Build a pipeline with mxbai-colbert-large-v1

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free