Listen to Motion: Robustly Learning Correlated Audio-Visual Representations
Zehan Wang, Xize Cheng, Li Tang, Luping Liu, Yang Zhao, Tao Jin, Chengfei Cai, WANG HongFa, Wei Liu, Zhou Zhao
OpenReview ground truth
Abstract
Audio-visual correlation learning has many applications and is pivotal in broader multimodal understanding and generation. Recently, many existing methods try to learn audio-visual contrastive representations from web-scale videos and show impressive performance. However, these methods mainly focus on learning the correlation between audio and static visual information (such as objects and background) while ignoring the crucial role of motion information in determining sounds in videos. Besides, the widespread presence of false and multiple positive audio-visual pairs in web-scale unlabeled videos also limits the performance of audio-visual representations. In this paper, we propose \textbf{Li}sten to \textbf{Mo}tion (LiMo) to capture motion information explicitly and align motion and audio robustly. Specifically, for modeling the motion in video, we extract the temporal visual semantic by facilitating the interaction between frames, while retaining static visual-audio correlation knowledge acquired in previous models. To prompt a more robust audio-visual alignment, we propose learning motion-audio alignment more specifically by distinguishing different clips within the same video. And we quantitatively measure the likelihood of each sample being false positive or containing multiple positive instances, then adaptively reweight samples in the final learning objective. Our extensive experiments demonstrate the effectiveness of LiMo on various audio-visual downstream tasks. On audio-visual retrieval, LiMo achieves absolute improvements of at least 15\% top1 accuracy on AudioSet and VGGSound. On our newly proposed motion-specific tasks, LiMo exhibits much better performance. Moreover, LiMo also achieves advanced accuracy on audio event recognition, demonstrating enhanced discriminability of audio representations.
Author context
Most prolific author: 12 submissions (credibility 0.79).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 57% of matchups.
- ▲ beat A Simple Romance Between Multi-Exit Vision… ×6
- ▼ lost to Tool-Augmented Reward Modeling ×6
- ▲ beat Exploring High-Order Message-Passing in Gr… ×4
- ▲ beat SPFormer: Enhancing Vision Transformer wit… ×4
- ▲ beat Acoustic Prompt Tuning: Empowering Large L… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)