NEWVectors or files. Pick a path.Start →
    Models/Document Analysis/PaddlePaddle/PP-DocLayoutV3
    HFDocument StructureApache 2.0

    PP-DocLayoutV3

    by PaddlePaddle

    High-accuracy document layout analysis with instance segmentation

    Identifiers
    Model ID
    PaddlePaddle/PP-DocLayoutV3
    Feature URI
    mixpeek://document_extractor@v1/paddle_pp_doclayoutv3_v1

    Overview

    PP-DocLayoutV3 is PaddlePaddle's third-generation document layout analysis model that combines object detection with instance segmentation for precise document structure understanding. Built on an efficient backbone with 33.3M parameters, it identifies and segments 23 document element types including text blocks, tables, figures, headers, footers, and mathematical formulas. The model uses a multi-scale feature pyramid network for handling elements of varying sizes.

    Architecture

    Detection + instance segmentation architecture built on PaddleDetection. Uses an FPN backbone for multi-scale feature extraction, with separate detection and segmentation heads. The model predicts bounding boxes and pixel-level masks for 23 document element categories simultaneously, enabling precise layout parsing even with overlapping or nested elements.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so PP-DocLayoutV3 runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // Boxes, masks, depth maps and anomaly scores are structured
              // results, not vectors. They go in payload and are reachable
              // through pre_filters on a retriever, not through similarity.
              payload: {
                detections: modelOutput,
                source_key: "archive/2026/asset-00412",
              },
            },
          ],
        }),
      },
    );
    
    // No managed alternative for an open label set. Two extractors do emit a
    // bbox, for the one thing each detects: document_graph_extractor@v1 per
    // layout block, face_identity_extractor@v1 per face. Nothing ships that
    // returns masks, depth maps or anomaly scores.

    Capabilities

    • Document layout detection
    • Instance segmentation of page elements
    • Table region detection
    • Figure and caption extraction
    • Mathematical formula localization

    Use Cases on Mixpeek

    Preprocessing documents before OCR for structured extraction
    Identifying table regions for specialized table parsers
    Segmenting scientific papers into sections and figures
    Converting scanned documents to structured digital formats

    Benchmarks

    DatasetMetricScoreSource
    PubLayNet[email protected]96.2Model card
    DocLayNet[email protected]79.8Model card
    CDLA[email protected]90.1Model card

    Performance

    Input SizeVariable
    GPU Latency~28ms per page on A100
    GPU Throughput~35 pages/sec
    GPU MemoryModel dependent

    Specification

    FrameworkHF
    OrganizationPaddlePaddle
    FeatureDocument Structure
    Outputstructure tokens
    Modalitiesdocument
    RetrieverSection Filter
    Parameters33.3M
    LicenseApache 2.0
    Downloads/mo400K

    Research Paper

    Model paper or technical report

    arxiv.org

    Build a pipeline with PP-DocLayoutV3

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free