PapersWithELO
← ICLR 2024 leaderboard

Listen to Motion: Robustly Learning Correlated Audio-Visual Representations

Zehan Wang, Xize Cheng, Li Tang, Luping Liu, Yang Zhao, Tao Jin, Chengfei Cai, WANG HongFa, Wei Liu, Zhou Zhao

representation learningaudio-visual representation learning
85.00100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
91.20100
Mimo
band ≈ ±21 pct pts (from σ = 0.42)
74.50100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.39)

OpenReview ground truth

Rejected

Abstract

Audio-visual correlation learning has many applications and is pivotal in broader multimodal understanding and generation. Recently, many existing methods try to learn audio-visual contrastive representations from web-scale videos and show impressive performance. However, these methods mainly focus on learning the correlation between audio and static visual information (such as objects and background) while ignoring the crucial role of motion information in determining sounds in videos. Besides, the widespread presence of false and multiple positive audio-visual pairs in web-scale unlabeled videos also limits the performance of audio-visual representations. In this paper, we propose \textbf{Li}sten to \textbf{Mo}tion (LiMo) to capture motion information explicitly and align motion and audio robustly. Specifically, for modeling the motion in video, we extract the temporal visual semantic by facilitating the interaction between frames, while retaining static visual-audio correlation knowledge acquired in previous models. To prompt a more robust audio-visual alignment, we propose learning motion-audio alignment more specifically by distinguishing different clips within the same video. And we quantitatively measure the likelihood of each sample being false positive or containing multiple positive instances, then adaptively reweight samples in the final learning objective. Our extensive experiments demonstrate the effectiveness of LiMo on various audio-visual downstream tasks. On audio-visual retrieval, LiMo achieves absolute improvements of at least 15\% top1 accuracy on AudioSet and VGGSound. On our newly proposed motion-specific tasks, LiMo exhibits much better performance. Moreover, LiMo also achieves advanced accuracy on audio event recognition, demonstrating enhanced discriminability of audio representations.

Author context

Most prolific author: 12 submissions (credibility 0.79).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 38 comparisons

Ranked above opponent in 57% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)