The Power of the Senses: Generalizable Manipulation from Vision and Touch through Masked Multimodal Learning
Carmelo Sferrazza, Younggyo Seo, Hao Liu, Youngwoon Lee, Pieter Abbeel
OpenReview ground truth
Abstract
Humans rely on the synergy of their senses for most essential tasks. For tasks requiring object manipulation, we seamlessly and effectively exploit the complementarity of our senses of vision and touch. This paper draws inspiration from such capabilities and aims to find a systematic approach to fuse visual and tactile information in a reinforcement learning setting. We propose Masked Multimodal Learning (M3L), which jointly learns a policy and visual-tactile representations based on masked autoencoding. The representations jointly learned from vision and touch improve sample efficiency, and unlock generalization capabilities beyond those achievable through each of the senses separately. Remarkably, representations learned in a multimodal setting also benefit vision-only policies at test time. We consider simulations provided of both visual and tactile observations, namely, a robotic insertion environment, a door opening task, and dexterous in-hand manipulation, demonstrating the benefits of learning a multimodal policy. Videos of the experiments are available at https://m3l.site. Code will be released upon acceptance.
Author context
Most prolific author: 17 submissions (credibility 0.43).
Delta if applied: -0.5 percentile
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 44 comparisons
Ranked above opponent in 52% of matchups.
- ▼ lost to Zero-Shot Robotic Manipulation with Pre-Tr… ×6
- ▼ lost to Tree-Planner: Efficient Close-loop Task Pl… ×6
- ▲ beat In-Depth Comparison of Regularization Meth… ×4
- ▲ beat RedMotion: Motion Prediction via Redundanc… ×4
- ▲ beat MindAgent: Emergent Gaming Interaction ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 44)