PapersWithELO
← ICLR 2024 leaderboard

The Implicit Bias of Stochastic AdaGrad-Norm on Separable Data

Ruinan Jin, Wei Liu, Baoxiang Wang

optimizationAdaGrad-NormLast-iterate convergenceStochastic optimizationImplicit Bias
69.10100
Fused
band ≈ ±15 pct pts (from σ = 0.31)
66.40100
Mimo
band ≈ ±22 pct pts (from σ = 0.44)
69.80100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.43)

OpenReview ground truth

Rejected

TL;DR — A theoretical paper about the implicit bias of AdaGrad-Norm

Abstract

This paper explores stochastic adaptive gradient descent, i.e., stochastic AdaGrad-Norm, with applications to linearly separable data sets. For the stochastic AdaGrad-Norm equipped with a wide range of sampling noise, we demonstrate its almost surely convergence result to the $\mathcal{L}^{2}$ max-margin solution. This means that stochastic AdaGrad-Norm has an implicit bias that yields good generalization, even without regularization terms. We show that the convergence rate of the direction is $o({1}/{\ln^{\frac{1-\epsilon}{2}}n})$. Our approach takes a novel stance by explicitly characterizing the $\mathcal{L}^{2}$ max-margin direction. By doing so, we overcome the challenge that arises from the dependency between the stepsize and the gradient, and also address the limitations in the traditional AdaGrad-Norm analysis.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 55% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)