← ICLR 2024 leaderboard

PASTA: Pretrained Action-State Transformer Agents

Raphael Boige, Yannis Flet-Berliac, Arthur Flajolet, Guillaume Richard, Thomas PIERROT

reinforcement learningself-supervised pre-trainingtransformer models
37.30100
Fused
band ≈ ±15 pct pts (from σ = 0.31)
37.70100
Mimo
band ≈ ±21 pct pts (from σ = 0.42)
43.40100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.45)

OpenReview ground truth

Rejected

TL;DR — We investigate different pre-training strategies for transformers in RL and show that a single multi-domain model pre-trained with a first-principles objective can be efficiently finetuned for increased performance in multiple RL downstream tasks.

Abstract

Self-supervised learning has brought about a revolutionary paradigm shift in various computing domains, including NLP, vision, and biology. Recent approaches involve pre-training transformer models on vast amounts of unlabeled data, serving as a starting point for efficiently solving downstream tasks. In the realm of reinforcement learning, researchers have recently adapted these approaches by developing models pre-trained on expert trajectories, enabling them to address a wide range of tasks, from robotics to recommendation systems. However, existing methods mostly rely on intricate pre-training objectives tailored to specific downstream applications. This paper presents a comprehensive investigation of models we refer to as pre-trained action-state transformer agents (PASTA). Our study uses a unified methodology and covers an extensive set of general downstream tasks including behavioral cloning, offline RL, sensor failure robustness, and dynamics change adaptation. Our goal is to systematically compare various design choices and provide valuable insights to practitioners for building robust models. Key highlights of our study include tokenization at the action and state component level, using fundamental pre-training objectives like next token prediction, training models across diverse domains simultaneously, and using parameter efficient fine-tuning (PEFT). The developed models in our study contain fewer than 10 million parameters and the application of PEFT enables fine-tuning of fewer than 10,000 parameters during downstream adaptation, allowing a broad community to use these models and reproduce our experiments. We hope that this study will encourage further research into the use of transformers with first-principles design choices to represent RL trajectories and contribute to robust policy learning.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 30 comparisons

Ranked above opponent in 48% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 30)