Enhancing Fine-Tuning Performance of Large-Scale Text-to-Image Models on Specialized Datasets
Yan-Lin Zhu, Peipei Yang
OpenReview ground truth
Abstract
Fine-tuning pre-trained large-scale text-to-image models on specialized datasets has gained popularity for downstream image generation tasks. However, direct fine-tuning of Stable-Diffusion on such datasets often falls short of yielding satisfactory outcomes. To delve into the underlying reasons, we introduce a novel perspective to investigate the intrinsic factors impacting fine-tuning outcomes. We identified that the limitations in fine-tuning stem from an inability to effectively improve text-image alignment and reduce text-image alignment drift. To tackle this issue, we leverage the powerful optimization capabilities of contrastive learning for feature distribution. By explicitly refining text feature representations during generation, we aim to enhance text-image alignment and minimize the alignment drift, thereby improving the fine-tuning performance on specialized datasets. Our approach is versatile, resource-efficient, and seamlessly integrates with existing controllable generation methods. Experimental results demonstrate a significant enhancement in fine-tuning performance achieved by our method.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 33% of matchups.
- ▼ lost to Constructing Sparse Neural Architecture wi… ×6
- ▲ beat Memoria: Hebbian Memory Architecture for H… ×6
- ▲ beat Regulating the level of manipulation in te… ×6
- ▼ lost to Generative Marginalization Models ×4
- ▼ lost to Molecule Relaxation by Reverse Diffusion w… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)