Embracing Diversity: Zero-shot Classification Beyond a Single Vector per Class
Mazda Moayeri, Michael Rabbat, Mark Ibrahim, Diane Bouchacourt
OpenReview ground truth
TL;DR — Improving zero-shot classification by leveraging under-utilized capabilities of VLMs to infer and explicitly account for diversity within classes
Abstract
Vision-language models for the first time enable open-world classification of objects without the need for any retraining. While this zero-shot paradigm marks a significant advance, even today’s best models exhibit skewed performance when objects are dissimilar from their typical depiction. Real world objects such as pears appear in a variety of forms --- from diced to whole, on a table or in a bowl --- yet standard VLM classifiers map all instances of a class to a single vector based on the class label. We argue that to represent this rich diversity within a class, zero-shot classification should move beyond a single vector. We propose a method to encode and account for diversity within a class using inferred attributes, still in the zero-shot setting without retraining. We find our method consistently outperforms standard zero-shot classification over a large suite of datasets encompassing hierarchies, diverse object states, and real-world geographic diversity. We also find our method scales efficiently to a large number of attributes to account for diversity---leading to more accurate predictions for atypical instances. Finally, we highlight how our method offers fine-grained human-interpretable explanations of model predictions. We hope this work spurs further research into the promise of zero-shot classification beyond a single class vector for capturing diversity in the world.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 54% of matchups.
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)