LoRA ensembles for large language model fine-tuning
Xi Wang, Laurence Aitchison, Maja Rudolph
OpenReview ground truth
Abstract
Finetuned LLMs often exhibit poor uncertainty quantification, manifesting as overconfidence, poor calibration, and unreliable prediction results on test data or out-of-distribution samples. One approach commonly used in vision for alleviating this issue is a deep ensemble, which constructs an ensemble by training the same model multiple times using different random initializations. However, there is a huge challenge to ensembling LLMs: the most effective LLMs are very, very large. Keeping a single LLM in memory is already challenging enough: keeping an ensemble of e.g. 5 LLMs in memory is impossible in many settings. To address these issues, we propose an ensemble approach using Low-Rank Adapters (LoRA), a parameter-efficient fine-tuning technique. Critically, these low-rank adapters represent a very small number of parameters, orders of magnitude less than the underlying pre-trained model. Thus, it is possible to construct large ensembles of LoRA adapters with almost the same computational overhead as using the original model. We find that LoRA ensembles, applied on its own or on top of pre-existing regularization techniques, gives consistent improvements in predictive accuracy and uncertainty quantification.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 39% of matchups.
- ▲ beat R-EDL: Relaxing Nonessential Settings of E… ×8
- ▼ lost to Statistical Inference for Deep Learning vi… ×4
- ▼ lost to Smoothing for exponential family dynamical… ×4
- ▼ lost to Escaping the Sample Trap: Fast and Accurat… ×4
- ▼ lost to Learning Forward Compatible Representation… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)