Improving the efficiency of conformal predictors via test-time augmentation
Divya M Shanmugam, Helen Lu, Swami Sankaranarayanan, John Guttag
OpenReview ground truth
TL;DR — Combining learned test-time augmentation policies with conformal predictors produces prediction sets that are up to 30% smaller, with no loss of coverage.
Abstract
In conformal classification, the goal is to output a _set_ of predicted classes, accompanied by a probabilistic guarantee that the set includes the true class. Conformal approaches have gained widespread traction across domains because they can be composed with existing classifiers to generate predictions with probabilistically valid uncertainty estimates. In practice, however, the utility of conformal prediction is limited by its tendency to yield large prediction sets. We study this phenomenon and provide insights into why large set sizes persist, even for conformal methods designed to produce small sets. Using these insights, we propose a method to reduce prediction set size while maintaining coverage. We use test-time augmentation to replace a classifier's predicted probabilities with probabilites aggregated over a set of augmentations. Our approach is flexible, computationally efficient, and effective. It can be combined with any conformal score, requires no model retraining, and reduces prediction set sizes by up to 30\%. We conduct an evaluation of the approach spanning three datasets, three models, two established conformal scoring methods, and multiple coverage values to show when and why test-time augmentation is a useful addition to the conformal pipeline.
Author context
Most prolific author: 4 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 49% of matchups.
- ▲ beat Duality of Information Flow: Insights in G… ×4
- ▼ lost to Faster Maximum Inner Product Search in Hig… ×4
- ▲ beat Learning Successor Representations with Di… ×4
- ▼ lost to FIITED: Fine-grained embedding dimension o… ×4
- ▲ beat Tube Loss: A Novel Approach for High Quali… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)