PapersWithELO
← ICLR 2024 leaderboard

UpFusion: Novel View Diffusion from Unposed Sparse View Observations

Bharath Raj Nagoor Kani, Hsin-Ying Lee, Sergey Tulyakov, Shubham Tulsiani

generative modelsNovel View SynthesisDiffusion3DGenerative ModelsTransformers
72.30100
Fused
band ≈ ±15 pct pts (from σ = 0.29)
71.60100
Mimo
band ≈ ±22 pct pts (from σ = 0.43)
69.70100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

TL;DR — A system that can perform novel view synthesis of an object given a sparse set of reference images without corresponding pose information.

Abstract

We propose UpFusion, a system that can perform novel view synthesis and infer 3D representations for an object given a sparse set of reference images without corresponding pose information. Current sparse-view 3D inference methods typically rely on camera poses to geometrically aggregate information from input views, but are not robust in-the-wild when such information is unavailable/inaccurate. In contrast, UpFusion sidesteps this requirement by learning to implicitly leverage the available images as context in a conditional generative model for synthesizing novel views. We incorporate two complementary forms of conditioning into diffusion models for leveraging the input views: a) via inferring query-view aligned features using a scene-level transformer, b) via intermediate attentional layers that can directly observe the input image tokens. We show that this mechanism allows generating high-fidelity novel views while improving the synthesis quality given additional (unposed) images. We evaluate our approach on the Co3D dataset and demonstrate the benefits of our method over pose-reliant alternates, Finally, we also show that our learned model can generalize beyond the training categories, and hope that this provides a stepping stone to reconstructing generic objects from in-the-wild image collections.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 51% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)