PapersWithELO
← ICLR 2024 leaderboard

Interpreting Adaptive Gradient Methods by Parameter Scaling for Learning-Rate-Free Optimization

Min-Kook Suh, Seung-Woo Seo

optimizationlearning-rate-free learningadaptive gradient methods
6.00100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
5.30100
Mimo
band ≈ ±22 pct pts (from σ = 0.43)
8.80100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.38)

OpenReview ground truth

Rejected

Abstract

We address the challenge of estimating the learning rate for adaptive gradient methods used in training deep neural networks. While several learning-rate-free approaches have been proposed, they are typically tailored for steepest descent. However, although steepest descent methods offer an intuitive approach to finding minima, many deep learning applications require adaptive gradient methods to achieve faster convergence. In this paper, we interpret adaptive gradient methods as steepest descent applied on parameter-scaled networks, proposing learning-rate-free adaptive gradient methods. Experimental results verify the effectiveness of this approach, demonstrating comparable performance to hand-tuned learning rates across various scenarios. This work extends the applicability of learning-rate-free methods, enhancing training with adaptive gradient methods.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 29% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)