NEWVectors or files. Pick a path.Start →
    Models/Document Analysis/nvidia/nemotron-page-elements-v3
    PyTorchDocument StructureNVIDIA Open Model License

    nemotron-page-elements-v3

    by nvidia

    Lightweight document layout detector for tables, charts, and text regions

    Identifiers
    Model ID
    nvidia/nemotron-page-elements-v3
    Feature URI
    mixpeek://document_extractor@v1/nvidia_nemotron_page_elements_v3

    Overview

    Nemotron Page Elements v3 is NVIDIA's compact 54M parameter document layout detection model built on the YOLOX architecture. It identifies six categories of document elements -- tables, charts, infographics, titles, text blocks, and headers/footers -- on 1024x1024 input images. Purpose-built for enterprise document RAG pipelines where layout detection is a preprocessing step before OCR or structured extraction.

    Architecture

    YOLOX anchor-free object detector with a DarkNet53 backbone and Feature Pyramid Network (FPN). Processes document page images at 1024x1024 resolution and outputs bounding boxes with class labels for six element types. The anchor-free design simplifies deployment and improves detection of elements with unusual aspect ratios (wide tables, tall infographics).

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so nemotron-page-elements-v3 runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // Boxes, masks, depth maps and anomaly scores are structured
              // results, not vectors. They go in payload and are reachable
              // through pre_filters on a retriever, not through similarity.
              payload: {
                detections: modelOutput,
                source_key: "archive/2026/asset-00412",
              },
            },
          ],
        }),
      },
    );
    
    // No managed alternative for an open label set. Two extractors do emit a
    // bbox, for the one thing each detects: document_graph_extractor@v1 per
    // layout block, face_identity_extractor@v1 per face. Nothing ships that
    // returns masks, depth maps or anomaly scores.

    Capabilities

    • Table region detection
    • Chart and infographic localization
    • Title and header identification
    • Text block segmentation
    • Page layout analysis

    Use Cases on Mixpeek

    Preprocessing documents before OCR to route regions to specialized extractors
    Identifying table regions for dedicated table parsers
    Filtering out headers and footers before text extraction
    Building document structure trees for RAG chunking strategies

    Benchmarks

    DatasetMetricScoreSource
    DocLayNet[email protected]78.4Model card
    PubLayNet[email protected]94.1Model card
    Internal (6-class)[email protected]91.7Model card

    Performance

    Input SizeVariable
    GPU Latency~15ms per page on A100
    GPU Throughput~65 pages/sec
    GPU MemoryModel dependent

    Specification

    FrameworkPyTorch
    Organizationnvidia
    FeatureDocument Structure
    Outputstructure tokens
    Modalitiesdocument
    RetrieverSection Filter
    Parameters54M
    LicenseNVIDIA Open Model License
    Downloads/moN/A

    Research Paper

    Model paper or technical report

    arxiv.org

    Build a pipeline with nemotron-page-elements-v3

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free