PapersWithELO
← ICLR 2024 leaderboard

Towards Pareto-Optimality for Test-Time Adaptation

JoonHo Jang, DongHyeok Shin, Byeonghu Na, HeeSun Bae, Il-chul Moon

transfer & meta learningTest-Time AdaptationPareto-OptimalitySharpness-Aware Minimization
26.50100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
23.50100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
38.80100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.38)

OpenReview ground truth

Rejected

TL;DR — We propose a new approach to update the model parameters toward Pareto-Optimality across all individual objectives in Test-Time Adaptation.

Abstract

Test-Time Adaptation (TTA) has been effective for mitigating the distribution shifts of test datasets by adapting a pre-trained model. The existing TTA approaches update the model parameters online toward the gradient descent direction by averaging individual objectives in the current batch. The averaged gradient can be biased by only a few instances in the batch, leading to conflict among individual objectives when updating the model. To prevent a negative effect from the gradient conflict among test instances, a model could have been updated by the gradient that is agreeable by all objectives in the batch. Therefore, we propose a new approach to update the model parameters toward Pareto-Optimality across all individual objectives in TTA. Particularly, this paper suggests an extended version of the Pareto optimization to anticipate unexpected distribution shifts during testing time. This extension is enabled by merging the sharpness-aware minimization into the Pareto optimization. We demonstrate the effectiveness of the proposed approaches through experiments on three benchmark datasets: CIFAR10-to-CIFAR10C, CIFAR100-to-CIFAR100C, and ImageNet-to-ImageNetC.

Author context

Most prolific author: 6 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)