PapersWithELO
← ICLR 2024 leaderboard

Expressive Modeling is Insufficient for Offline RL: A Tractable Inference Perspective

Xuejie Liu, Anji Liu, Guy Van den Broeck, Yitao Liang

reinforcement learningOffline Reinforcement LearningTractable Probabilistic Models
87.70100
Fused
band ≈ ±16 pct pts (from σ = 0.32)
87.20100
Mimo
band ≈ ±23 pct pts (from σ = 0.46)
88.50100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.43)

OpenReview ground truth

Rejected

Abstract

A popular paradigm for offline Reinforcement Learning (RL) tasks is to first fit the offline trajectories to a sequence model, and then prompt the model for actions that lead to high expected return. While a common consensus is that more expressive sequence models imply better performance, this paper highlights that tractability, the ability to exactly and efficiently answer various probabilistic queries, plays an equally important role. Specifically, due to the fundamental stochasticity from the offline data-collection policies and the environment dynamics, highly non-trivial conditional/constrained generation is required to elicit rewarding actions. While it is still possible to approximate such queries, we observe that such crude estimates significantly undermine the benefits brought by expressive sequence models. To overcome this problem, this paper proposes Trifle (Tractable Inference for Offline RL), which leverages modern Tractable Probabilistic Models (TPMs) to bridge the gap between good sequence models and high expected returns at evaluation time. Empirically, Trifle achieves the most state-of-the-art scores in 9 Gym-MuJoCo benchmarks against strong baselines. Further, owing to its tractability, Trifle significantly outperforms prior approaches in stochastic environments and safe RL tasks (i.e. with state/action constraints) with minimum algorithmic modifications.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 30)