Diffusion Models for Open-Vocabulary Segmentation
Laurynas Karazija, Iro Laina, Andrea Vedaldi, Christian Rupprecht
OpenReview ground truth
TL;DR — We leverage a pre-trained diffusion model to perform open-vocabulary semantic segmentation by sampling support images for feature correlation without further training.
Abstract
The variety of objects in the real world is unlimited and is thus impossible to capture using models trained on a closed, pre-defined set of categories. Recently, open-vocabulary recognition has garnered significant attention, largely facilitated by advances in large-scale vision-language modelling. In this paper, we present OVDiff, a novel method that leverages the generative properties of text-to-image diffusion models for open-vocabulary segmentation. Specifically, we propose to synthesise support image sets from arbitrary textual categories, creating for each category a set of prototypes representative of both the category itself and its surrounding context (background). Our method relies solely on pre-trained components: segmentation is obtained by simply comparing a target image to the prototypes without further fine-tuning. We show that our method can be used to ground any pre-trained self-supervised feature extractor in natural language and provide explainable predictions by mapping back to regions in the support set. Our approach shows strong performance on a range of open-vocabulary segmentation benchmarks, obtaining a lead of more than 10% over prior work on PASCAL VOC.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 61% of matchups.
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)