NEWVectors or files. Pick a path.Start →
    Models/Text Extraction/PaddlePaddle/PaddleOCR-VL-1.5
    HFOCRapache-2.0

    PaddleOCR-VL-1.5

    by PaddlePaddle

    Vision-language OCR handling text, tables, formulas, charts, and 109 languages in 0.9B params

    15Kdl/month
    661likes
    959Mparams
    Identifiers
    Model ID
    PaddlePaddle/PaddleOCR-VL-1.5
    Feature URI
    mixpeek://image_extractor@v1/paddle_ocr_vl_15_v1

    Overview

    PaddleOCR-VL 1.5 replaces the traditional OCR pipeline (detect → recognize → layout) with a single vision-language model that understands document structure natively. At 0.9B parameters, it handles text recognition, table extraction, formula parsing, chart understanding, seal detection, text spotting, and 109 languages, including rare scripts like Tibetan and Bengali.

    On Mixpeek, PaddleOCR-VL replaces brittle multi-stage OCR pipelines with a single model call that produces structured output from any document type. Its robustness to scanning artifacts, skew, and poor lighting makes it reliable for real-world document ingestion.

    Architecture

    Vision-language model with document-specific pretraining. 0.9B parameters. Unified multi-task architecture handles text detection, recognition, layout analysis, table extraction, and formula parsing in a single forward pass. Robust to image degradation (scanning, warping, screen capture).

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so PaddleOCR-VL-1.5 runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The model produces text, so it lands in payload. Give the
              // collection a text vector index and embed that text to make it
              // searchable rather than only filterable.
              payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
              vectors: { "text-embedding": embeddingOfModelOutput },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // universal_extractor@v1 runs google/gemini-embedding-2
    // (3072-d) over a bucket, with no inference of your own.

    Capabilities

    • 109 language support including rare scripts
    • Unified text, table, formula, chart, and seal parsing
    • SOTA robustness to scanning artifacts, skew, and lighting
    • 0.9B parameters: 3-4x smaller than competing VLM-OCR models
    • Apache 2.0 license

    Use Cases on Mixpeek

    Document digitization: extract structured text from scanned archives
    Multilingual document processing: handle mixed-language business documents
    Form extraction: parse handwritten and printed forms into structured data
    Invoice and receipt processing: extract line items, totals, and metadata

    Benchmarks

    DatasetMetricScoreSource
    OmniDocBench v1.5Overall94.5%PaddlePaddle, 2026: arxiv,2601.21957
    Real5-OmniDocBench (robustness)OverallSOTAPaddlePaddle, 2026: arxiv,2601.21957

    Performance

    Input SizeDocument page image (variable)
    GPU Latency~35ms / page (A100)
    GPU Throughput~28 pages/sec (A100)
    GPU Memory~2 GB

    Specification

    FrameworkHF
    OrganizationPaddlePaddle
    FeatureOCR
    Outputtext + bbox
    Modalitiesvideo, image, document
    RetrieverText-in-Image
    Parameters959M
    Licenseapache-2.0
    Downloads/mo15K
    Likes661

    Research Paper

    PaddleOCR-VL-1.5: Multi-Task 0.9B VLM for Robust Document Parsing

    arxiv.org

    Build a pipeline with PaddleOCR-VL-1.5

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free