PapersWithELO
← ICLR 2024 leaderboard

CLIP as Multi-Task Multi-Kernel Learning

Tianjun Ke, Yucong Lin, Xingpeng Xia, Jiaheng Yin, Jiaxing Xu, Tianxi Cai, Junwei Lu

metric & kernel learningContrastive Language-Image PretrainingReproducing Kernel Hilbert SpaceMulti-Task Multi-Kernel Learning
42.30100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
44.40100
Mimo
band ≈ ±20 pct pts (from σ = 0.39)
40.50100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

Abstract

Contrastive Language-Image Pretraining (CLIP) is a foundational model that learns a latent embedding space through an inner product-based objective. In this paper, we provide a theoretical interpretation of CLIP utilizing Reproducing Kernel Hilbert Space (RKHS) framework. Specifically, we reformulate the problem of estimating the infinite-dimensional mapping with a neural network as selecting an unknown RKHS using multiple kernel learning. Such connection motivates us to propose to estimate the CLIP embedding via the multi-task multi-kernel (MTMK) method: we reformulate the different labels in the CLIP training data as the multiple training tasks, and reformulate learning the unknown CLIP embedding as choosing an optimal kernel from a family of Reproducing Kernel Hilbert Spaces, which is computationally more efficient. Utilizing the MTMK interpretation of CLIP, we also show an optimal statistical rate of the MTMK classifier under the scenario that both the number of covariates and the number of candidate kernels can increase with the sample size. Besides the synthetic simulations, we apply the proposed method to align the medical imaging data with the clinical codes in electronic health records and illustrate that our approach can learn the proper kernel space aligning the imaging embedding with the text embeddings with high accuracy.

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 38 comparisons

Ranked above opponent in 50% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)