PapersWithELO
← ICLR 2024 leaderboard

Visual Category Discovery via Linguistic Anchoring

Sua Choi, Dahyun Kang, Minsu Cho

self/semi-supervised learningmulti-modalclusteringcategory discovery
35.40100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
38.50100
Mimo
band ≈ ±21 pct pts (from σ = 0.41)
37.60100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.43)

OpenReview ground truth

Rejected

Abstract

We address the problem of generalized category discovery (GCD) that aims to classify entire images of a partially labeled image collection with the total number of target classes being unknown. Motivated by the relevance of visual category to linguistic semantics, we propose language-anchored contrastive learning for GCD. Assuming consistent relations between images and their corresponding texts in an image-text joint embedding space, our method incorporates image-text consistency constraints into contrastive learning. To perform this process without manual image-text annotations, we assign each image with a corresponding text embedding by retrieving $k$-nearest-neighbor words among a random corpus of diverse words and aggregating them through cross-attention. The proposed method achieves state-of-the-art performance on the standard benchmarks, ImageNet100, CUB, Stanford Cars, and Herbarium19.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)