← ICLR 2024 leaderboard

Boosting Semi-Supervised Learning via Variational Confidence Calibration and Unlabeled Sample Elimination

Qianhan Feng, Shijie Fang, Tong Lin

self/semi-supervised learningSemi-Supervised LearningCalibrationSample Elimination
21.30100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
21.10100
Mimo
band ≈ ±21 pct pts (from σ = 0.43)
24.10100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.42)

OpenReview ground truth

Rejected

Abstract

Despite the recent progress of Semi-supervised Learning (SSL), we argue that the existing methods may not employ unlabeled examples effectively and efficiently. Many pseudo-label-based methods select unlabeled examples into the training stage based on the inaccurate confidence scores provided by the output layer of the classifier network. Additionally, most prior work typically adpots all the available unlabeled examples without data pruning, which is incapable of learning from massive unlabeled data. To address these issues, this paper proposes two methods called VCC (Variational Confidence Calibration) and INFUSE (INfluence-Function-based Unlabeled Sample Elimination). VCC is a general-purpose plugin of confidence calibration for SSL. By approximating the calibrated confidence through three types of consistency scores, a variational autoencoder is leveraged to reconstruct the confidence score for selecting more accurate pseudo-labels. Based on the influence function, INFUSE is a data pruning method for constructing a core dataset of unlabeled examples. The effectiveness of our methods is demonstrated through experiments on multiple datasets and in various settings. For example, on the CIFAR-100 dataset with 400 labeled examples, VCC reduces the classification error rate of FixMatch from 46.47\% to 43.31\% (with improvement of 3.16\%). On the SVHN dataset with 250 labeled examples, INFUSE achieves 2.61\% error rate using only 10\% unlabeled data, which is better than RETRIEVE (2.90\%) and the baseline with full unlabeled data (3.80\%). Putting all the pieces together, the combined VCC-INFUSE plugins can reduce the error rate of FlexMatch from 26.49\% to 25.41\% on the CIFAR100 dataset (with improvement of 1.08\%) while saving nearly half of the original training time (from 223.96 GPU hours to 115.47 GPU hours).

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)