Offline Imitation Learning without Auxiliary High-quality Behavior Data
Jie-Jing Shao, Hao-Sen Shi, Tian Xu, Lan-Zhe Guo, Yang Yu, Yu-Feng Li
OpenReview ground truth
Abstract
In this work, we study the problem of Offline Imitation Learning (OIL), where an agent aims to learn from the demonstrations composed of expert behaviors and sub-optimal behaviors without additional online environment interactions. Previous studies typically assume that there is high-quality behavioral data mixed in the auxiliary offline data and seriously degrades when only low-quality data from an off-policy distribution is available. In this work, we break through the bottleneck of OIL relying on auxiliary high-quality behavior data and make the first attempt to demonstrate that low-quality data is also helpful for OIL. Specifically, we utilize the transition information from offline data to maximize the policy transition probability towards expert-observed states. This guidance can improve long-term returns on states that are not observed by experts when reward signals are not available, ultimately enabling imitation learning to benefit from low-quality data. We instantiate our proposition in a simple but effective algorithm, Behavioral Cloning with Dynamic Programming (BCDP), which involves executing behavioral cloning on the expert data and dynamic programming on the unlabeled offline data respectively. In the experiments on benchmark tasks, unlike most existing offline imitation learning methods that do not utilize low-quality data sufficiently, our BCDP algorithm can still achieve an average performance gain of more than 40\% even when the offline data is purely random exploration.
Author context
Most prolific author: 9 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 57% of matchups.
- ▼ lost to Tree Search-Based Policy Optimization unde… ×6
- ▼ lost to TransCues: Boundary and Reflection-empower… ×6
- ▼ lost to Optimal Sample Complexity for Average Rewa… ×4
- ▼ lost to Achieving Minimax Optimal Sample Complexit… ×4
- ▲ beat One is More: Diverse Perspectives within a… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)