← ICLR 2024 leaderboard

Is the Glass Half-Empty or Half-Full? A Mixture-Of-Tasks Perspective on Missing Modality

Daniel Yang, Tiantian Feng, Yoonsoo Nam, Jihwan Lee, Shrikanth Narayanan

representation learningmissing modalitymodality competitionmultimodal learningmultimodal fusion
37.00100
Fused
band ≈ ±15 pct pts (from σ = 0.29)
29.80100
Mimo
band ≈ ±21 pct pts (from σ = 0.42)
45.10100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

TL;DR — We propose the Missing Modality Performance Testbed (MMPT) to tackle issues with missing modality in multimodal learning setups by reconstructing missing modality robustness analysis as a fundamental part of multimodal representation learning.

Abstract

A common issue with multimodal learning setups is the unavailability of one or more modalities. Historically, missing modality has been treated as a matter of robustness, aiming to prevent performance degradation caused by stochastic loss of training and testing modalities. However, this perspective does not align with many scientific and industrial use cases of deep models where unimodal inputs are more common than having multiple modalities. Moreover, it poses practical challenges such as complicating comparisons between studies and causing ambiguity in understanding optimal model behavior. We instead propose a `glass-half-full' approach---the Missing Modality Performance Testbed (MMPT)--- which sheds light on the pivotal elements for enhancing model performance under the effect of missing modalities. MMPT reconceptualizes missing modality robustness analysis as a fundamental aspect of multimodal representation learning. This formulation allows us to connect missing modality to modality competition, an area of work that aims to improve unimodal representations in a multimodal context for late-fusion models. We create a unified framework for both missing modality and modality competition by relaxing their architectural assumptions. Via this linkage, we explore how current approaches to missing modality impact the underlying model representations and the requisite representations for favorable performance. We validate this novel perspective on a wide variety of multimodal datasets with the intention of enabling simple and clear benchmarking for future research. Finally, we present a new state-of-the-art in missing modality performance and identify potential areas for further improvement.

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 49% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)