BiRefNet
by ZhengPeng7
High-resolution foreground segmentation for object masks and visual evidence cleanup
ZhengPeng7/BiRefNetmixpeek://image_extractor@v1/zhengpeng7_birefnet_v1Overview
BiRefNet is the official checkpoint for Bilateral Reference for High-Resolution Dichotomous Image Segmentation. It targets foreground/background masks, salient object segmentation, and related cases where the useful evidence is an object region rather than the whole image.
On Mixpeek, BiRefNet can turn images or sampled video frames into mask metadata. Agents can use those masks to filter frames with clear foreground objects, crop objects before embedding, or remove distracting background before downstream OCR, detection, captioning, or similarity search.
Architecture
Image-segmentation model for high-resolution dichotomous segmentation. The Hugging Face card lists Transformers support through AutoModelForImageSegmentation, MIT licensing, and tags for background removal, mask generation, camouflaged object detection, and salient object detection.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so BiRefNet runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// Boxes, masks, depth maps and anomaly scores are structured
// results, not vectors. They go in payload and are reachable
// through pre_filters on a retriever, not through similarity.
payload: {
detections: modelOutput,
source_key: "archive/2026/asset-00412",
},
},
],
}),
},
);
// No managed alternative for an open label set. Two extractors do emit a
// bbox, for the one thing each detects: document_graph_extractor@v1 per
// layout block, face_identity_extractor@v1 per face. Nothing ships that
// returns masks, depth maps or anomaly scores.Capabilities
- Foreground/background mask generation
- High-resolution dichotomous image segmentation
- Background removal and object isolation
- Useful pre-processing for embeddings, OCR, and VLM captioning
- MIT license
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| Hugging Face | Monthly downloads | 824K | HF model metadata, June 2026 |
| BiRefNet task coverage | Segmentation tags | DIS, camouflaged, salient object | BiRefNet model card |
Performance
Run before visual embeddings when foreground isolation improves retrieval quality
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
Bilateral Reference for High-Resolution Dichotomous Image Segmentation
arxiv.orgBuild a pipeline with BiRefNet
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free