UpFusion: Novel View Diffusion from Unposed Sparse View Observations
Bharath Raj Nagoor Kani, Hsin-Ying Lee, Sergey Tulyakov, Shubham Tulsiani
OpenReview ground truth
TL;DR — A system that can perform novel view synthesis of an object given a sparse set of reference images without corresponding pose information.
Abstract
We propose UpFusion, a system that can perform novel view synthesis and infer 3D representations for an object given a sparse set of reference images without corresponding pose information. Current sparse-view 3D inference methods typically rely on camera poses to geometrically aggregate information from input views, but are not robust in-the-wild when such information is unavailable/inaccurate. In contrast, UpFusion sidesteps this requirement by learning to implicitly leverage the available images as context in a conditional generative model for synthesizing novel views. We incorporate two complementary forms of conditioning into diffusion models for leveraging the input views: a) via inferring query-view aligned features using a scene-level transformer, b) via intermediate attentional layers that can directly observe the input image tokens. We show that this mechanism allows generating high-fidelity novel views while improving the synthesis quality given additional (unposed) images. We evaluate our approach on the Co3D dataset and demonstrate the benefits of our method over pose-reliant alternates, Finally, we also show that our learned model can generalize beyond the training categories, and hope that this provides a stepping stone to reconstructing generic objects from in-the-wild image collections.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 51% of matchups.
- ▲ beat Enhancing Fine-Tuning Performance of Large… ×4
- ▼ lost to Generative Marginalization Models ×4
- ▲ beat S\(^{2}\)-DMs: Skip-Step Diffusion Models ×4
- ▲ beat Tree Search-Based Policy Optimization unde… ×4
- ▼ lost to SEPT: Towards Efficient Scene Representati… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)