PapersWithELO
← ICLR 2024 leaderboard

Universal Guidance for Diffusion Models

Arpit Bansal, Hong-Min Chu, Avi Schwarzschild, Roni Sengupta, Micah Goldblum, Jonas Geiping, Tom Goldstein

generative modelsGenerative ModelsComputer VisionDiffusion Models
79.30100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
58.00100
Mimo
band ≈ ±21 pct pts (from σ = 0.43)
93.00100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.41)

OpenReview ground truth

Accepted

TL;DR — We propose universal guidance, an algorithm that achieves conditional generation with any base diffusion model and guidance functions without any retraining.

Abstract

Typical diffusion models are trained to accept a particular form of conditioning, most commonly text, and cannot be conditioned on other modalities without retraining. In this work, we propose a universal guidance algorithm that enables diffusion models to be controlled by arbitrary guidance modalities without the need to retrain any use-specific components. We show that our algorithm successfully generates quality images with guidance functions including segmentation, face recognition, object detection, style guidance and classifier signals.

Author context

Most prolific author: 11 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)