Spade : Training-Free Improvement of Spatial Fidelity in Text-to-Image Generation
Agneet Chatterjee, Yiran Luo, Chitta Baral, Yezhou Yang
OpenReview ground truth
TL;DR — Improving Spatial Understanding of Text to Image Models using Spatially Accurate Reference Images in a Training-free manner
Abstract
Text-to-Image (T2I) generation models have seen progressive improvements in their abilities to generate photo-realistic images. However, it has been demonstrated that they struggle to follow reasoning-intensive textual instructions, particularly when it comes to generating accurate spatial relationships between objects. In this work, we present an approach to improve upon the above shortcomings of these models by leveraging spatially accurate images (LSAI) as grounding reference to guide diffusion-based T2I models. Given an input prompt containing a spatial phrase, our method involves symbolically creating a corresponding synthetic image, which accurately represents the spatial relationship articulated in the prompt. Next, we use the created image alongside the text prompt, in a training-free manner to condition image synthesis models in generating spatially coherent images. To facilitate our LSAI method, we create SPADE, a large database of 190k text-image pairs, where each image is deterministically generated through open-source 3D rendering tools encompassing a diverse set of 80 MS-COCO objects. Variation of the images in SPADE is introduced through object and background manipulation as well as GPT-4 guided layout arrangement. We evaluate our method of utilizing SPADE as T2I guidance on Stable Diffusion and ControlNet, and find our LSAI method substantially improves upon existing methods on the VISOR benchmark. Through extensive ablations and analysis, we analyze LSAI with respect to multiple facets of SPADE and also perform human studies to demonstrate the effectiveness of our method on prompts which contain multiple relationships and out-of-distribution objects. Finally, we present our SPADE Generator as an extendable framework to the research community, emphasizing its potential for expansion.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 48% of matchups.
- ▼ lost to SE(3)-Stochastic Flow Matching for Protein… ×4
- ▲ beat Beyond Vanilla Variational Autoencoders: D… ×4
- ▼ lost to GOAt: Explaining Graph Neural Networks via… ×4
- ▼ lost to Smoothing for exponential family dynamical… ×4
- ▲ beat Can Synthetic Data Reduce Conservatism of … ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)