CLIP-Guided Reinforcement Learning for Open-Vocabulary Tasks
Haobin Jiang, Zongqing Lu
OpenReview ground truth
Abstract
Open-vocabulary ability is crucial for an agent designed to follow natural language instructions. In this paper, we focus on developing an open-vocabulary agent through reinforcement learning. We leverage the capability of CLIP to segment the target object specified in language instructions from the image observations. The resulting confidence map replaces the text instruction as input to the agent's policy, grounding the natural language into the visual information. Compared to the giant embedding space of natural language, the two-dimensional confidence map provides a more accessible unified representation for neural networks. When faced with instructions containing unseen objects, the agent converts textual descriptions into comprehensible confidence maps as input, enabling it to accomplish open-vocabulary tasks. Additionally, we introduce an intrinsic reward function based on the confidence map to more effectively guide the agent towards the target objects. Our single-task experiments demonstrate that our intrinsic reward significantly improves performance. In multi-task experiments, through testing on tasks out of the training set, we show that the agent, when provided with confidence maps as input, possesses open-vocabulary capabilities.
Author context
Most prolific author: 11 submissions (credibility 0.64).
Delta if applied: -0.1 percentile
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 40% of matchups.
- ▲ beat Contrastive Graph Autoencoder for Geometri… ×6
- ▼ lost to NL2ProGPT: Taming Large Language Model for… ×6
- ▼ lost to Optimal Sample Complexity for Average Rewa… ×4
- ▼ lost to Efficient Action Robust Reinforcement Lear… ×4
- ▼ lost to Revisiting the Static Model in Robust Rein… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)