PapersWithELO
← ICLR 2024 leaderboard

Offline Imitation Learning without Auxiliary High-quality Behavior Data

Jie-Jing Shao, Hao-Sen Shi, Tian Xu, Lan-Zhe Guo, Yang Yu, Yu-Feng Li

reinforcement learningimitation learningoffline imitation learningoffline reinforcement learning
77.00100
Fused
band ≈ ±15 pct pts (from σ = 0.31)
77.50100
Mimo
band ≈ ±23 pct pts (from σ = 0.46)
73.10100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.42)

OpenReview ground truth

Rejected

Abstract

In this work, we study the problem of Offline Imitation Learning (OIL), where an agent aims to learn from the demonstrations composed of expert behaviors and sub-optimal behaviors without additional online environment interactions. Previous studies typically assume that there is high-quality behavioral data mixed in the auxiliary offline data and seriously degrades when only low-quality data from an off-policy distribution is available. In this work, we break through the bottleneck of OIL relying on auxiliary high-quality behavior data and make the first attempt to demonstrate that low-quality data is also helpful for OIL. Specifically, we utilize the transition information from offline data to maximize the policy transition probability towards expert-observed states. This guidance can improve long-term returns on states that are not observed by experts when reward signals are not available, ultimately enabling imitation learning to benefit from low-quality data. We instantiate our proposition in a simple but effective algorithm, Behavioral Cloning with Dynamic Programming (BCDP), which involves executing behavioral cloning on the expert data and dynamic programming on the unlabeled offline data respectively. In the experiments on benchmark tasks, unlike most existing offline imitation learning methods that do not utilize low-quality data sufficiently, our BCDP algorithm can still achieve an average performance gain of more than 40\% even when the offline data is purely random exploration.

Author context

Most prolific author: 9 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 32 comparisons

Ranked above opponent in 57% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)