Video Generation Beyond a Single Clip
Hsin-Ping Huang, Yu-Chuan Su, Ming-Hsuan Yang
OpenReview ground truth
Abstract
This work tackles the challenge of generating long videos, which entails producing videos that surpass the output length of video generation models. Due to computational constraints, video generation models are restricted to generating relatively short video clips compared to the length of real-world videos. Existing approaches employ a sliding window technique to generate long videos during inference, but this method is often restricted to homogeneous content and recurring events. To generate long videos that encompass diverse content and multiple events, we propose utilizing additional guidance to steer the video generation process. We further introduce a multi-stage approach to address this challenge, enabling us to leverage existing video generation models to produce high-quality videos within a limited time window while holistically modeling the long video based on the provided guidance. Our method complements existing video generation efforts. Extensive experiments on challenging real-world videos demonstrate the advantages of the proposed method, which outperforms the state-of-the-art by up to 9.5\% in objective metrics and is preferred by users over 80\% of the time. The source code and trained models will be released to the public.
Author context
Most prolific author: 7 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 37% of matchups.
- ▲ beat DOG: Discriminator-only Generation Beats G… ×6
- ▼ lost to From Images to Connections: Can DQN with G… ×6
- ▼ lost to Denoising Diffusion Bridge Models ×4
- ▼ lost to GAIA: Zero-shot Talking Avatar Generation ×4
- ▼ lost to Closed-Form Diffusion Models ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)