NEWVectors or files. Pick a path.Start →
    Models/Text Extraction/datalab-to/chandra-ocr-2
    HFOCROpenRAIL-M

    chandra-ocr-2

    by datalab-to

    High-accuracy multilingual OCR with 90+ language support

    Identifiers
    Model ID
    datalab-to/chandra-ocr-2
    Feature URI
    mixpeek://image_extractor@v1/datalab_chandra_ocr2_v1

    Overview

    Chandra OCR 2 from Datalab is a 5B parameter vision-language model optimized for optical character recognition across 90+ languages. Built on the Qwen3.5 architecture, it achieves 85.9% on the olmOCR benchmark (SOTA at time of release) and 77.8% multilingual accuracy across 43 languages, a 12-point improvement over v1. The model outputs structured Markdown, HTML, or JSON and excels at handwriting, tables, and mathematical formulas.

    Architecture

    Image-text-to-text model based on Qwen3.5 with a vision encoder fine-tuned for document understanding. Processes full-page document images and generates structured text output (Markdown with table formatting and LaTeX math). The dual-encoder architecture handles both the visual layout and text content simultaneously.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so chandra-ocr-2 runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The model produces text, so it lands in payload. Give the
              // collection a text vector index and embed that text to make it
              // searchable rather than only filterable.
              payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
              vectors: { "image-embedding": embeddingOfModelOutput },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // image_extractor@v1 runs google/siglip-base-patch16-224
    // (768-d) over a bucket, with no inference of your own.

    Capabilities

    • 90+ language OCR
    • Handwriting recognition
    • Table structure extraction
    • Mathematical formula recognition
    • Markdown/HTML/JSON output

    Use Cases on Mixpeek

    Digitizing multilingual document archives
    Extracting structured data from scanned forms and invoices
    Processing handwritten notes and medical records
    Converting academic papers with equations to searchable text

    Benchmarks

    DatasetMetricScoreSource
    olmOCRAccuracy85.9%Model card
    Multilingual (43 langs)Accuracy77.8%Model card
    Table RecognitionTEDS92.1%Model card

    Performance

    Input SizeVariable
    GPU Latency~200ms per page on A100
    GPU Throughput~5 pages/sec
    GPU MemoryModel dependent

    Specification

    FrameworkHF
    Organizationdatalab-to
    FeatureOCR
    Outputtext + bbox
    Modalitiesvideo, image, document
    RetrieverText-in-Image
    Parameters5.3B
    LicenseOpenRAIL-M
    Downloads/mo1.45M

    Research Paper

    Model paper or technical report

    arxiv.org

    Build a pipeline with chandra-ocr-2

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free