Uncertainty-Aware Decision Transformer for Stochastic Driving Environments
Zenan Li, Fan Nie, Qiao Sun, Fang Da, Hang Zhao
OpenReview ground truth
TL;DR — We present an uncertainty-aware decision transformer to maximize cumulative rewards at certain states but plan cautiously at states with stochastic environment transitions.
Abstract
Offline Reinforcement Learning (RL) has emerged as a promising framework for learning policies without active interactions, making it especially appealing for autonomous driving tasks. Recent successes of Transformers inspire casting offline RL as sequence modeling, which performs well in long-horizon tasks. However, they are overly optimistic in stochastic environments with incorrect assumptions that the same goal can be consistently achieved by identical actions. In this paper, we introduce an uncertainty-aware decision transformer (UNREST) for planning in stochastic driving environments without introducing additional transition or complex generative models. Specifically, UNREST estimates state uncertainties by the conditional mutual information between transitions and returns, and segments sequences accordingly. Discovering the 'uncertainty accumulation' and 'temporal locality' properties of driving environments, UNREST replaces the global returns in decision transformers with less uncertain truncated returns, to learn from true outcomes of agent actions rather than environment transitions. We also dynamically evaluate environmental uncertainty during inference for cautious planning. Extensive experimental results demonstrate UNREST's superior performance in various driving scenarios and the power of our uncertainty estimation strategy.
Author context
Most prolific author: 9 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 42 comparisons
Ranked above opponent in 45% of matchups.
- ▼ lost to Variational Language Concepts for Interpre… ×6
- ▼ lost to ZeroFlow: Scalable Scene Flow via Distilla… ×4
- ▲ beat In-Depth Comparison of Regularization Meth… ×4
- ▲ beat Memory-efficient particle filter recurrent… ×4
- ▲ beat RedMotion: Motion Prediction via Redundanc… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 42)