On the Joint Interaction of Models, Data, and Features
Yiding Jiang, Christina Baek, J Zico Kolter
OpenReview ground truth
TL;DR — We propose a framework for feature learning that can explain previously not understood phenommena in deep learning.
Abstract
Learning features from data is one of the defining characteristics of deep learning, but the theoretical understanding of the role features play in deep learning is still in early development. To address this gap, we introduce a new tool, the interaction tensor, for empirically analyzing the interaction between data and model through features. With the interaction tensor, we make several key observations about how features are distributed in data and how models with different random seeds learn different features. Based on these observations, we propose a conceptual framework for feature learning. Under this framework, the expected accuracy for a single hypothesis and agreement for a pair of hypotheses can both be derived in closed form. We demonstrate that the proposed framework can explain empirically observed phenomena, including the recently discovered Generalization Disagreement Equality (GDE) that allows for estimating the generalization error with only unlabeled data. Further, our theory also provides explicit construction of natural data distributions that break the GDE. Thus, we believe this work provides valuable new insight into our understanding of feature learning.
Author context
Most prolific author: 18 submissions (credibility 0.32).
Delta if applied: -0.7 percentile
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 51% of matchups.
- ▲ beat Understanding Deep Neural Networks as Dyna… ×4
- ▲ beat Information based explanation methods for … ×4
- ▲ beat Conditional MAE: An Empirical Study of Mul… ×4
- ▼ lost to Efficient Integrators for Diffusion Genera… ×4
- ▼ lost to The Curse of Diversity in Ensemble-Based E… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)