Universal Guidance for Diffusion Models
Arpit Bansal, Hong-Min Chu, Avi Schwarzschild, Roni Sengupta, Micah Goldblum, Jonas Geiping, Tom Goldstein
OpenReview ground truth
TL;DR — We propose universal guidance, an algorithm that achieves conditional generation with any base diffusion model and guidance functions without any retraining.
Abstract
Typical diffusion models are trained to accept a particular form of conditioning, most commonly text, and cannot be conditioned on other modalities without retraining. In this work, we propose a universal guidance algorithm that enables diffusion models to be controlled by arbitrary guidance modalities without the need to retrain any use-specific components. We show that our algorithm successfully generates quality images with guidance functions including segmentation, face recognition, object detection, style guidance and classifier signals.
Author context
Most prolific author: 11 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 56% of matchups.
- ▲ beat GAIA: Zero-shot Talking Avatar Generation ×6
- ▲ beat Optimisation-Based Multi-Modal Semantic Im… ×4
- ▲ beat LOVECon: Text-driven Training-free Long Vi… ×4
- ▲ beat Probabilistic Graphical Model for Robust G… ×4
- ▼ lost to Scaling Laws of RoPE-based Extrapolation ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)