PapersWithELO
← ICLR 2024 leaderboard

Detect Every Thing with Few Examples

Xinyu Zhang, Yuting Wang, Abdeslam Boularias

representation learningFew-shotObject detectionOpen-vocabulary
94.80100
Fused
band ≈ ±17 pct pts (from σ = 0.33)
90.40100
Mimo
band ≈ ±23 pct pts (from σ = 0.47)
95.50100
DeepSeek
band ≈ ±24 pct pts (from σ = 0.47)

OpenReview ground truth

Rejected

TL;DR — We propose an open-set object detector that uses example images rather than language as category representation, outperforming state-of-the-art in open-vocabulary, few-shot, one-shot object detection benchmarks.

Abstract

Open-set object detection aims at detecting arbitrary categories beyond those seen during training. Most recent advancements have adopted the open-vocabulary paradigm, utilizing vision-language backbones to represent categories with language. In this paper, we introduce DE-ViT, an open-set object detector that employs vision-only DINOv2 backbones and learns new categories through example images instead of language. To improve general detection ability, we transform multi-classification tasks into binary classification tasks while bypassing per-class inference, and propose a novel region propagation technique for localization. We evaluate DE-ViT on open-vocabulary, few-shot, and one-shot object detection benchmark with COCO and LVIS. For COCO, DE-ViT outperforms the open-vocabulary SoTA by 6.9 AP50 and achieves 50 AP50 in novel classes. DE-ViT surpasses the few-shot SoTA by 15 mAP on 10-shot and 7.2 mAP on 30-shot and one-shot SoTA by 2.8 AP50. For LVIS, DE-ViT outperforms the open-vocabulary SoTA by 2.2 mask AP and reaches 34.3 mask APr.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)