On the Evaluation of Generative Models in Distributed Learning Tasks
Zixiao Wang, Farzan Farnia, Zhenghao Lin, Yunheng Shen, Bei Yu
OpenReview ground truth
TL;DR — FID score can lead to inconsistent rankings of generative models in heterogeneous distributed learning settings, while KID score cannot.
Abstract
The evaluation of deep generative models including generative adversarial networks (GANs) and diffusion models has been extensively studied in the literature. While the existing evaluation methods mainly target a centralized learning problem with training data stored by a single client, many applications of generative models concern distributed learning settings, e.g. the federated learning scenario, where training data are collected by and distributed among several clients. In this paper, we study the evaluation of generative models in distributed learning tasks with heterogeneous data distributions. First, we focus on the Fréchet inception distance (FID) and consider the following FID-based aggregate scores over the clients: 1) FID-avg as the mean of clients' individual FID scores, 2) FID-all as the FID distance of the trained model to the collective dataset containing all clients' data. We prove that the model rankings according to the FID-all and FID-avg scores could be inconsistent, which can lead to different optimal generative models according to the two aggregate scores. Next, we consider the kernel inception distance (KID) and similarly define the KID-avg and KID-all aggregations. Unlike the FID case, we prove that KID-all and KID-avg result in the same rankings of generative models. We perform several numerical experiments on standard image datasets and training schemes to support our theoretical findings on the evaluation of generative models in distributed learning problems.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 53% of matchups.
- ▼ lost to Learning Energy-Based Models by Cooperativ… ×6
- ▲ beat AutoHall: Automated Hallucination Dataset … ×4
- ▲ beat Tell, Don't Show: Internalized Reasoning i… ×4
- ▼ lost to Unifying Feature and Cost Aggregation with… ×4
- ▲ beat Potential Based Diffusion Motion Planning ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)