PapersWithELO
← ICLR 2024 leaderboard

MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation

Mingzhen Sun, Weining Wang, Yanyuan Qiao, Longteng Guo, Jiahui Sun, Xinxin Zhu, Jing Liu

generative modelsSounding video generationmulti-modal generationdiffusion model
80.90100
Fused
band ≈ ±16 pct pts (from σ = 0.31)
81.50100
Mimo
band ≈ ±23 pct pts (from σ = 0.45)
71.10100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.44)

OpenReview ground truth

Rejected

Abstract

Sounding video generation (SVG) is a challenging audio-video joint generation task that requires both single-modal realism and cross-modal consistency. Previous diffusion-based methods tackled SVG within the original signal space, resulting in a huge computation burden. In this paper, we introduce a novel multi-modal latent diffusion model (MM-LDM), which establishes a perceptual latent space that is perceptually equivalent to the original audio-video signal space but drastically reduces computational complexity. We unify the representation of audio and video signals and construct a shared high-level semantic feature space to bridge the information gap between audio and video modalities. Furthermore, we present a novel cross-modal sampling guidance that extends our generative models to audio-to-video and video-to-audio conditional generation tasks. We obtain the new state-of-the-art results with significant quality and efficiency gains. In particular, our method achieves an overall improvement in all evaluation metrics and a faster training and sampling speed.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 30)