Optimisation-Based Multi-Modal Semantic Image Editing
Bowen Li, Yongxin Yang, Steven McDonagh, Shifeng Zhang, Petru-Daniel Tudosiu, Sarah Parisot
OpenReview ground truth
TL;DR — We propose an inference-time optimisation editing method using diffusion models, designed to extend beyond textual edits to accommodate multiple instruction types (e.g., spatial layout-based; pose, scribbles, edge maps).
Abstract
Image editing affords increased control over the aesthetics and content of generated images. Pre-existing works focus predominantly on text-based instructions to achieve desired image modifications, which limit edit precision and accuracy. In this work, we propose an inference-time editing optimisation, designed to extend beyond textual edits to accommodate multiple editing instruction types (e.g., spatial layout-based; pose, scribbles, edge maps). We propose to disentangle the editing task into two competing subtasks: successful local image modifications and global content consistency preservation, where subtasks are guidedthrough two dedicated loss functions. By allowing to adjust the influence of each loss function, we build a flexible editing solution that can be adjusted to user preferences. We evaluate our method using text, pose and scribble edit conditions, and highlight our ability to achieve complex edits, through both qualitative and quantitative experiments.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 40 comparisons
Ranked above opponent in 33% of matchups.
- ▲ beat A Data-Driven Measure of Relative Uncertai… ×12
- ▲ beat Forward Explanation : Why Catastrophic For… ×12
- ▲ beat Simple mechanisms for representing, indexi… ×10
- ▼ lost to Metanetwork: A novel approach to interpret… ×8
- ▼ lost to RoBERT: Low-Cost Bi-Directional Sequence M… ×8
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 40)