PapersWithELO
← ICLR 2024 leaderboard

Enhancing Fine-Tuning Performance of Large-Scale Text-to-Image Models on Specialized Datasets

Yan-Lin Zhu, Peipei Yang

generative modelsfine-tuning pre-trained modelsstable-diffusiondiffusion modelscontrastive learning
9.30100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
7.50100
Mimo
band ≈ ±20 pct pts (from σ = 0.39)
9.30100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

Abstract

Fine-tuning pre-trained large-scale text-to-image models on specialized datasets has gained popularity for downstream image generation tasks. However, direct fine-tuning of Stable-Diffusion on such datasets often falls short of yielding satisfactory outcomes. To delve into the underlying reasons, we introduce a novel perspective to investigate the intrinsic factors impacting fine-tuning outcomes. We identified that the limitations in fine-tuning stem from an inability to effectively improve text-image alignment and reduce text-image alignment drift. To tackle this issue, we leverage the powerful optimization capabilities of contrastive learning for feature distribution. By explicitly refining text feature representations during generation, we aim to enhance text-image alignment and minimize the alignment drift, thereby improving the fine-tuning performance on specialized datasets. Our approach is versatile, resource-efficient, and seamlessly integrates with existing controllable generation methods. Experimental results demonstrate a significant enhancement in fine-tuning performance achieved by our method.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 33% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)