PATHS: Parameter-wise Adaptive Two-Stage Training Harnessing Scene Transition Mask Adapters for Video Retrieval
SeongMin Kang, Yoon-Sik Cho
OpenReview ground truth
Abstract
Image-text pre-trained model, e.g., CLIP, has gained significant traction even in the field of video-text learning. Recent approaches extended CLIP to video tasks, and have achieved unprecedented performances in the foundational study of video understanding: text-video retrieval. However, unlike conventional transfer learning within the same domain, transfer learning across different modalities from images to videos often requires fine-tuning the whole pre-trained weights rather than keeping them frozen. This may result in overfitting and distorting the pre-trained weights, leading to a degradation in performance. To address this challenge, we introduce a learning strategy, termed Parameter-wise Adaptive Two-stage training Harnessing Scene transition mask adapter (PATHS). Our two-stage learning process alleviates the deviations of the pre-trained weights. A novel method of finding the optimal weights is used in the first stage, which efficiently narrows down to strong candidates by only monitoring the fluctuations of parameters. Once the parameters are fixed to optimal values, the second stage is dedicated to acquiring knowledge of scenes with an adapter module. PATHS can be applied to any existing models in a plug-and-play manner, and always achieves performance improvements from the base models. We report state-of-the-art performances across key text-video benchmark datasets, including MSRVTT and LSMDC. Our code is available at https://anonymous.4open.science/r/PATHS_.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 34% of matchups.
- ▼ lost to Demystifying CLIP Data ×4
- ▲ beat Heterogeneity of Regularization between ad… ×4
- ▼ lost to Zero-Level-Set Encoder for Neural Distance… ×4
- ▲ beat Motion PointNet: Solving Dynamic Capture i… ×4
- ▲ beat Rethinking the Effectiveness of Graph Clas… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)