MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
Mingzhen Sun, Weining Wang, Yanyuan Qiao, Longteng Guo, Jiahui Sun, Xinxin Zhu, Jing Liu
OpenReview ground truth
Abstract
Sounding video generation (SVG) is a challenging audio-video joint generation task that requires both single-modal realism and cross-modal consistency. Previous diffusion-based methods tackled SVG within the original signal space, resulting in a huge computation burden. In this paper, we introduce a novel multi-modal latent diffusion model (MM-LDM), which establishes a perceptual latent space that is perceptually equivalent to the original audio-video signal space but drastically reduces computational complexity. We unify the representation of audio and video signals and construct a shared high-level semantic feature space to bridge the information gap between audio and video modalities. Furthermore, we present a novel cross-modal sampling guidance that extends our generative models to audio-to-video and video-to-audio conditional generation tasks. We obtain the new state-of-the-art results with significant quality and efficiency gains. In particular, our method achieves an overall improvement in all evaluation metrics and a faster training and sampling speed.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 30 comparisons
Ranked above opponent in 59% of matchups.
- ▲ beat Generative Pre-Trained Speech Language Mod… ×6
- ▼ lost to Language Model Beats Diffusion - Tokenizer… ×4
- ▲ beat Towards Minimal Targeted Updates of Langua… ×4
- ▼ lost to In Search of the Long-Tail: Systematic Gen… ×4
- ▲ beat Amplifying Training Data Exposure through … ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 30)