PapersWithELO
← ICLR 2024 leaderboard

Extending to New Domains without Visual and Textual Oracles

Daiqing Qi, Handong Zhao, Aidong Zhang, Sheng Li

representation learningmultimodal learningvision-language modeldomain adaptationdomain generalization
17.00100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
23.10100
Mimo
band ≈ ±21 pct pts (from σ = 0.42)
17.30100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.39)

OpenReview ground truth

Rejected

TL;DR — With the help of a pre-trained vision-language model, we extending a model to new domains without any information (both textual and visual information) of them

Abstract

To avoid the high cost of collecting visual data from all test domains in domain adaption task, recent work takes advantage of the pre-trained large-scale vision language models and augment training data with only text descriptions (e.g.,“a photo/painting/sketch...”) of each test domain. However, in many real-world ap- plications, such text information of test domains is not always available in ad- vance. Moreover, even if we can verbalize all test domains, it is laborious for existing work (Dunlap et al., 2023) to train a different augmentation network for each possible unseen domain. To overcome these challenges, we benefit from the multimodal embedding space of a pre-trained vision-language model and propose to acquire training-free and domain-invariant augmentations with text descrip- tions of arbitrary crafted unseen domains, which not necessarily match test do- mains. Beyond achieving state-of-the-art results, compared with existing works that require trainable augmentation networks, our approach is also notably more time-efficient, and exhibits a more solid theoretical support.

Author context

Most prolific author: 11 submissions (credibility 0.64).

Delta if applied: -0.1 percentile

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 38% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)