Adapting Cross-View Localization to New Areas without Ground Truth Positions
Zimin Xia, Yujiao Shi, Hongdong Li, Julian F. P. Kooij
OpenReview ground truth
Abstract
Given a ground-level query image, cross-view localization aims to estimate the location of the ground camera by matching the query to a geo-referenced aerial image that covers the local surroundings. Recent works have focused on developing powerful frameworks trained with ground truth (GT) locations of ground images within aerial images. However, the trained models always suffer a performance drop when applied to images in a new target area that differs from the training data. In most deployment scenarios, acquiring accurate GT location data for target-area images to re-train the network can be expensive and sometimes infeasible. In contrast, collecting images with coarse GT with errors of tens of meters is relatively easier. Motivated by this, our paper focuses on improving the generalization of a trained model by leveraging only the target area images without accurate GT. We propose a weakly-supervised learning approach based on knowledge self-distillation, namely, using predictions from a teacher model to supervise a student model with the same architecture. Our approach includes a mode-based pseudo GT generation for reducing uncertainty in pseudo GT and an outlier filtering to remove unreliable pseudo GT for student training. We validate our approach is generic by performing experiments on two recent state-of-the-art models with two benchmarks. The results demonstrate that our approach consistently and considerably boosts the localization performance in the target area.
Author context
Most prolific author: 4 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 30 comparisons
Ranked above opponent in 45% of matchups.
- ▼ lost to ZeroFlow: Scalable Scene Flow via Distilla… ×4
- ▲ beat Memory-efficient particle filter recurrent… ×4
- ▼ lost to Exploiting Implicit Rigidity Constraints v… ×4
- ▲ beat Backdoor Attack for Federated Learning wit… ×4
- ▲ beat CLIP Facial Expression Recognition: Balanc… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 30)