PapersWithELO
← ICLR 2024 leaderboard

RoboGPT : An intelligent agent of making embodied long-term decisions for daily instruction tasks

Yaran Chen, cuiwenbo2023@ia.ac.cn, chenyw220@163.com, tanmining@163.com, trisoil@bupt.edu.cn, Dongbin Zhao, He Wang

robotics & planningRobot task planningDaily tasks following by instructionsEmbodied AI
11.40100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
13.40100
Mimo
band ≈ ±19 pct pts (from σ = 0.37)
11.50100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

TL;DR — RoboGPT agent solves daily insturction task with long-term decisions through LLM planning and low-level policy

Abstract

Robotic agents must master common sense and long-term sequential decisions to solve daily tasks through natural language instruction. The developments in Large Language Models (LLMs) in natural language processing have inspired efforts to use LLMs in complex robot planning. Despite LLMs' great generalization and comprehension of instructional tasks, LLM-generated task plans sometimes lack feasibility and correctness. To address the problem, we propose a RoboGPT agent for making embodied long-term decisions for daily tasks, with two modules: 1) LLM-based planning with Re-Plan to break the task into multiple sub-goals; 2) RoboSkill individually designed for sub-goals to learn better navigation and manipulation skills. The LLM-based planning is enhanced with a new robotic dataset and re-plan, called RoboGPT. The new robotic dataset of 67k daily instruction tasks is gathered for fine-tuning the LLaMA model and obtaining RoboGPT. RoboGPT palnner with strong generalization can plan hundreds of daily tasks, and re-plan based on the environment, thereby addressing the nomenclature diversity challenge. Additionally, a low-computational Re-Plan module is designed to allow plans to flexibly adapt to the environment. The proposed RoboGPT agent outperforms SOTA methods on the ALFRED daily tasks. Moreover, RoboGPT palnner exceeds SOTA LLM-based planners like ChatGPT in task-planning rationality for hundreds of unseen daily tasks, and even other domain tasks, while keeping the large model's original broad application and generality.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 35% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)