PapersWithELO
← ICLR 2024 leaderboard

Lightweight Unsupervised Federated Learning with Pretrained Vision Language Model

Hao Yan, Yuhong Guo

fairness, safety & privacyunsupervised federated learningvision-language model
14.30100
Fused
band ≈ ±16 pct pts (from σ = 0.32)
11.50100
Mimo
band ≈ ±23 pct pts (from σ = 0.46)
13.40100
DeepSeek
band ≈ ±23 pct pts (from σ = 0.46)

OpenReview ground truth

Rejected

Abstract

Federated learning aims to tackle the ``isolated data island" problem, where it trains a collective model from physically isolated clients while safeguarding the privacy of users' data. However, supervised federated learning necessitates that each client labels their data for training, which can be both time-consuming and resource-intensive, and may even be impractical for edge devices. Moreover, the training and transmission of deep models present challenges to the computation and communication capabilities of the clients. To address these two inherent challenges in supervised federated learning, we propose a novel lightweight unsupervised federated learning approach that leverages unlabeled data on each client to perform lightweight model training and communication by harnessing pretrained vision-language models, such as CLIP. By capitalizing on the zero-shot prediction capability and the well-trained image encoder of the pre-trained CLIP model, we have carefully crafted an efficient and resilient self-training approach. This method refines the initial zero-shot predicted pseudo-labels of unlabeled instances through the sole training of a linear classifier on top of the fixed image encoder. Additionally, to address data heterogeneity within each client, we propose a class-balanced text feature sampling strategy for generating synthetic instances in the feature space to support local training. Experiments are conducted on multiple benchmark datasets. The experimental results demonstrate that our proposed method greatly enhances model performance in comparison to CLIP's zero-shot predictions and even outperforms supervised federated learning benchmark methods given limited computational and communication overhead.

Author context

Most prolific author: 4 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 28 comparisons

Ranked above opponent in 42% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 28)