PapersWithELO
← ICLR 2024 leaderboard

AugUndo: Scaling Up Augmentations for Unsupervised Depth Completion

Yangchao Wu, Tian Yu Liu, Hyoungseob Park, Stefano Soatto, Dong Lao, Alex Wong

representation learningData AugmentationMonocular Depth Completion
36.60100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
38.10100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
40.30100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.45)

OpenReview ground truth

Rejected

Abstract

Unsupervised depth completion methods are trained predominantly using structure-from-motion. The training objective involves minimizing photometric reconstruction error between temporally (from video) or spatially (from stereo) adjacent images, which assumes photometric consistency in co-visible regions across frames. Block artifacts from geometric transformations, intensity saturation, and occlusions are amongst the many undesirable by-products of common data augmentation schemes that affect reconstruction quality, and thus the resulting model performance. Hence, typical data augmentations on the image that are viewed as essential to training pipelines in other vision tasks have seen limited use beyond small image intensity changes and flipping. In fact, the sparse depth modality have seen even less variety as intensity transformations alter the scale of the measured 3D scene, and geometric transformations may decimate the sparse points during resampling. We propose a method that unlocks a wide range of previously-infeasible geometric augmentations for unsupervised depth completion. This is achieved by reversing,or ``undo"-ing, geometric transformations to the coordinates of the output depth, warping the depth map back to the original reference frame. This enables computing the photometric reprojection loss via the original images and sparse depth maps, eliminating the pitfalls resulting from naive loss computation on the augmented inputs. This simple yet effective strategy allows us to scale up augmentations to boost performance. We demonstrate our method on indoor (VOID) and outdoor (KITTI) datasets where we improve upon three existing methods by an average of 10.4% overall across both datasets.

Author context

Most prolific author: 7 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 32 comparisons

Ranked above opponent in 49% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)