Extending to New Domains without Visual and Textual Oracles
Daiqing Qi, Handong Zhao, Aidong Zhang, Sheng Li
OpenReview ground truth
TL;DR — With the help of a pre-trained vision-language model, we extending a model to new domains without any information (both textual and visual information) of them
Abstract
To avoid the high cost of collecting visual data from all test domains in domain adaption task, recent work takes advantage of the pre-trained large-scale vision language models and augment training data with only text descriptions (e.g.,“a photo/painting/sketch...”) of each test domain. However, in many real-world ap- plications, such text information of test domains is not always available in ad- vance. Moreover, even if we can verbalize all test domains, it is laborious for existing work (Dunlap et al., 2023) to train a different augmentation network for each possible unseen domain. To overcome these challenges, we benefit from the multimodal embedding space of a pre-trained vision-language model and propose to acquire training-free and domain-invariant augmentations with text descrip- tions of arbitrary crafted unseen domains, which not necessarily match test do- mains. Beyond achieving state-of-the-art results, compared with existing works that require trainable augmentation networks, our approach is also notably more time-efficient, and exhibits a more solid theoretical support.
Author context
Most prolific author: 11 submissions (credibility 0.64).
Delta if applied: -0.1 percentile
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 38% of matchups.
- ▲ beat OneBNet: Binarized Neural Networks using D… ×6
- ▼ lost to Tag2Text: Guiding Vision-Language Model vi… ×4
- ▼ lost to Motion PointNet: Solving Dynamic Capture i… ×4
- ▼ lost to OTMatch: Improving Semi-Supervised Learnin… ×4
- ▼ lost to Fully Hyperbolic Convolutional Neural Netw… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)