Visual Category Discovery via Linguistic Anchoring
Sua Choi, Dahyun Kang, Minsu Cho
OpenReview ground truth
Abstract
We address the problem of generalized category discovery (GCD) that aims to classify entire images of a partially labeled image collection with the total number of target classes being unknown. Motivated by the relevance of visual category to linguistic semantics, we propose language-anchored contrastive learning for GCD. Assuming consistent relations between images and their corresponding texts in an image-text joint embedding space, our method incorporates image-text consistency constraints into contrastive learning. To perform this process without manual image-text annotations, we assign each image with a corresponding text embedding by retrieving $k$-nearest-neighbor words among a random corpus of diverse words and aggregating them through cross-attention. The proposed method achieves state-of-the-art performance on the standard benchmarks, ImageNet100, CUB, Stanford Cars, and Herbarium19.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 46% of matchups.
- ▲ beat Forked Diffusion for Conditional Graph Gen… ×6
- ▼ lost to A Recipe for Improved Certifiable Robustne… ×4
- ▼ lost to On robust overfitting: adversarial trainin… ×4
- ▲ beat P4Q: Learning to Prompt for Quantization i… ×4
- ▲ beat Continual Nonlinear ICA-Based Representati… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)