PapersWithELO
← ICLR 2024 leaderboard

CLIP-Guided Reinforcement Learning for Open-Vocabulary Tasks

Haobin Jiang, Zongqing Lu

reinforcement learningopen-vocabulary
14.40100
Fused
band ≈ ±15 pct pts (from σ = 0.29)
12.60100
Mimo
band ≈ ±22 pct pts (from σ = 0.44)
13.20100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.39)

OpenReview ground truth

Rejected

Abstract

Open-vocabulary ability is crucial for an agent designed to follow natural language instructions. In this paper, we focus on developing an open-vocabulary agent through reinforcement learning. We leverage the capability of CLIP to segment the target object specified in language instructions from the image observations. The resulting confidence map replaces the text instruction as input to the agent's policy, grounding the natural language into the visual information. Compared to the giant embedding space of natural language, the two-dimensional confidence map provides a more accessible unified representation for neural networks. When faced with instructions containing unseen objects, the agent converts textual descriptions into comprehensible confidence maps as input, enabling it to accomplish open-vocabulary tasks. Additionally, we introduce an intrinsic reward function based on the confidence map to more effectively guide the agent towards the target objects. Our single-task experiments demonstrate that our intrinsic reward significantly improves performance. In multi-task experiments, through testing on tasks out of the training set, we show that the agent, when provided with confidence maps as input, possesses open-vocabulary capabilities.

Author context

Most prolific author: 11 submissions (credibility 0.64).

Delta if applied: -0.1 percentile

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 40% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)