PapersWithELO
← ICLR 2024 leaderboard

Test like you Train in Implicit Deep Learning

Zaccharie Ramzi, Pierre Ablin, Gabriel Peyré, Thomas Moreau

optimizationdeep equilibrium modelsimplicit differentiationbilevel optimizationbilevelimplicit deep learningmeta-learning
79.20100
Fused
band ≈ ±17 pct pts (from σ = 0.34)
81.20100
Mimo
band ≈ ±24 pct pts (from σ = 0.47)
74.20100
DeepSeek
band ≈ ±24 pct pts (from σ = 0.48)

OpenReview ground truth

Rejected

TL;DR — We show that Deep Equilibrium Models (DEQs) do not in practice benefit from a higher number of inner iterations at test-time compared to that used in training.

Abstract

Implicit deep learning has recently gained popularity with applications ranging from meta-learning to Deep Equilibrium Networks~(DEQs). In its very general formulation, it relies on expressing some components of deep learning pipelines implicitly, typically via a root equation called the inner problem. In practice, the solution of the inner problem is approximated with an iterative procedure, usually with a fixed number of inner iterations during training. At inference time, the inner problems needs to be solved with new data. A popular belief is that increasing the number of inner iterations relative to the one used in training yields better performances. In this paper, we question such an assumption and provide a detailed theoretical analysis in a simple affine setting. We demonstrate that overparametrization plays a key role: increasing the number of iterations at test time cannot improve performances for overparametrized networks. We validate our theory on an array of implicit deep-learning problems. We show that DEQs, which are typically overparametrized, do not benefit from increasing the number of iterations at inference while meta-learning, which is typically not overparametrized, benefits from it.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 28)