Zero-shot Inversion Process for Image Attribute Editing with Diffusion Models
Zhanbo Feng, Zenan Ling, Feng Zhou, Ci Gong, Jie LI, Robert C Qiu
OpenReview ground truth
Abstract
Denoising diffusion models have shown outstanding performance in image editing. Existing works tend to use either image-guided methods, which provide a visual reference but lack control over semantic coherence, or text-guided methods, which ensure faithfulness to text guidance but lack visual quality. To address the problem, we propose the Zero-shot Inversion Process (ZIP), a framework that injects a fusion of generated visual reference and text guidance into the semantic latent space of a frozen pre-trained diffusion model. Only using a tiny neural network, the proposed ZIP produces diverse content and attributes under the intuitive control of the text prompt. Moreover, ZIP shows remarkable robustness for both in-domain and out-of-domain attribute manipulation on real images. We perform detailed experiments on various benchmark datasets. Compared to state-of-the-art methods, ZIP produces images of equivalent quality while providing a realistic editing effect.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 35% of matchups.
- ▲ beat Regulating the level of manipulation in te… ×6
- ▲ beat I Know You Did Not Write That! A Sampling … ×6
- ▼ lost to The Blessing of Randomness: SDE Beats ODE … ×4
- ▼ lost to Efficient Integrators for Diffusion Genera… ×4
- ▼ lost to Towards More Accurate Diffusion Model Acce… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)