GeRA: Label-Efficient Geometrically Regularized Alignment
Dustin Klebe, Tal Shnitzer, Mikhail Yurochkin, Leonid Karlinsky, Justin Solomon
OpenReview ground truth
TL;DR — This research paper introduces a semi-supervised method to align multi-modal data distributions by using the geometric manifold structures of pretrained unimodal encoders.
Abstract
Pretrained unimodal encoders incorporate rich semantic information into embedding space structures. To be similarly informative, multi-modal encoders typically require massive amounts of paired data for alignment and training. We introduce a semi-supervised Geometrically Regularized Alignment (GeRA) method to align the embedding spaces of pretrained unimodal encoders in a label-efficient way. Our method leverages the manifold geometry of unpaired (unlabeled) data to improve alignment performance. To prevent distortions to local geometry during the alignment process —potentially disrupting semantic neighborhood structures and causing misalignment of unobserved pairs — we introduce a geometric loss term. This term is built upon a diffusion operator that captures the local manifold geometry of the unimodal pretrained encoders. GeRA is modality-agnostic and thus can be used to align pretrained encoders from any data modalities. We provide empirical evidence to the effectiveness of our method in the domains of speech-text and image-text alignment. Our experiments demonstrate significant improvement in alignment quality compared to a variaty of leading baselines, especially with a small amount of paired data, using our proposed geometric regularization.
Author context
Most prolific author: 7 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 45% of matchups.
- ▲ beat Barycentric Alignment of Mutually Disentan… ×4
- ▲ beat Hybrid Sharing for Multi-Label Image Class… ×4
- ▲ beat Musketeer: Joint Training/Inference for Mu… ×4
- ▲ beat Mixture of LoRA Experts ×4
- ▼ lost to TETA: Temporal-Enhanced Text-to-Audio Gene… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)