PapersWithELO
← ICLR 2024 leaderboard

RAND: Robustness Aware Norm Decay For Quantized Seq2seq Models

David Qiu, David Rim, Shaojin Ding, Oleg Rybakov, Yanzhang He

representation learningautomatic speech recognitionquantizationseq2seqtransducer
45.30100
Fused
band ≈ ±16 pct pts (from σ = 0.32)
38.70100
Mimo
band ≈ ±23 pct pts (from σ = 0.45)
48.80100
DeepSeek
band ≈ ±23 pct pts (from σ = 0.45)

OpenReview ground truth

Rejected

Abstract

With the rapid increase in the size of neural networks, model compression has become an important area of research. Quantization is an effective technique at decreasing the model size, memory access, and compute load of large models. Despite recent advances in quantization aware training (QAT) technique, most papers present evaluations that are focused on computer vision tasks, which have different layer composition and training dynamics compared to sequence tasks. In this paper, we first benchmark the impact of popular techniques such as straight through estimator, pseudo-quantization noise (PQN), learnable scale parameter, clipping, etc. on 4-bit seq2seq models across a suite of speech recognition datasets ranging from 1,000 hours to 1 million hours, as well as one machine translation dataset to illustrate its applicability outside of speech. Through the experiments, we report that accuracy suffers when there is insufficient regularization signal flowing back to the outliers. We propose to construct the quantization scale as different functions of the outliers in order to regularize them as part of the end-to-end learning problem (outperforming popular learnable scale and clipping methods). PQN-QAT shows a larger improvement under the proposed method, and it opens up the possibility to exploit some of its other benefits: 1) training a single model that performs well in mixed precision mode and 2) improved generalization on long form speech recognition.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 28)