Towards Aligned Layout Generation via Diffusion Model with Aesthetic Constraints
Jian Chen, Ruiyi Zhang, Yufan Zhou, Changyou Chen
OpenReview ground truth
TL;DR — Unified model for layout generation using constrained diffusion.
Abstract
Controllable layout generation refers to the process of creating a plausible visual arrangement of elements within a graphic design (*e.g.*, document and web designs) with constraints representing design intentions. Although recent diffusion-based models have achieved state-of-the-art FID scores, they tend to exhibit more pronounced misalignment compared to earlier transformer-based models. In this work, we propose the **LA**yout **C**onstraint diffusion mod**E**l (LACE), a unified model to handle a broad range of layout generation tasks, such as arranging elements with specified attributes and refining or completing a coarse layout design. The model is based on continuous diffusion models. Compared with existing methods that use discrete diffusion models, continuous state-space design can enable the incorporation of continuous aesthetic constraint functions in training more naturally. For conditional generation, we propose injecting layout conditions in the form of masks or gradient guidance during inference. Empirical results show that LACE produces high-quality layouts and outperforms existing state-of-the-art baselines. We will release our source code and model checkpoints.
Author context
Most prolific author: 7 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 30 comparisons
Ranked above opponent in 40% of matchups.
- ▲ beat Potential Based Diffusion Motion Planning ×6
- ▼ lost to Navigating the Design Space of Equivariant… ×4
- ▼ lost to DreamFlow: High-quality text-to-3D generat… ×4
- ▲ beat Generative Modeling with Phase Stochastic … ×4
- ▲ beat Large Scene Synthesis Controlled With Deta… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 30)