← ICLR 2024 leaderboard

Video Decomposition Prior: Editing Videos Layer by Layer

Gaurav Shrivastava, Ser-Nam Lim, Abhinav Shrivastava

self/semi-supervised learningVideo decompositionTest-time optimization techniqueVideo RelightingVideo DehazingUnsupervised Video Object SegmentationEdits Propagation
28.90100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
26.10100
Mimo
band ≈ ±19 pct pts (from σ = 0.37)
26.90100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.41)

OpenReview ground truth

Accepted

TL;DR — We introduce a novel inference-time optimization framework that performs three primary vision task of video relighting, unsupervised object segmentation and video dehazing leveraging the compositionality inherent in the videos.

Abstract

In the evolving landscape of video editing methodologies, a majority of deep learning techniques are often reliant on extensive datasets of observed input and ground truth sequence pairs for optimal performance. Such reliance often falters when acquiring data becomes challenging, especially in tasks like video dehazing and relighting, where replicating identical motions and camera angles in both corrupted and ground truth sequences is complicated. Moreover, these conventional methodologies perform best when the test distribution closely mirrors the training distribution. Recognizing these challenges, this paper introduces a novel video decomposition prior `VDP' framework which derives inspiration from professional video editing practices. Our methodology does not mandate task-specific external data corpus collection, instead pivots to utilizing the motion and appearance of the input video. VDP framework decomposes a video sequence into a set of multiple RGB layers and associated opacity levels. These set of layers are then manipulated individually to obtain the desired results. We addresses tasks such as video object segmentation, dehazing, and relighting. Moreover, we introduce a novel logarithmic video decomposition formulation for video relighting tasks, setting a new benchmark over the existing methodologies. We evaluate our approach on standard video datasets like DAVIS, REVIDE, & SDSD and show qualitative results on a diverse array of internet videos.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 40)