MindAgent: Emergent Gaming Interaction
Ran Gong, Qiuyuan Huang, Xiaojian Ma, Hoi Vo, Zane Durante, Yusuke Noda, Zilong Zheng, Demetri Terzopoulos, Li Fei-Fei, Jianfeng Gao
OpenReview ground truth
Abstract
Large Language Models (LLMs) can perform complex scheduling in a multi-agent system and can coordinate agents to complete sophisticated tasks that require extensive collaboration. However, despite the introduction of numerous gaming frameworks, the community lacks adequate benchmarks that support the implementation of a general multi-agent infrastructure encompassing collaboration between LLMs and human-NPCs. We propose a novel infrastructure--- MindAgent---for evaluating planning and coordination-emergent capabilities in the context of gaming interaction. In particular, our infrastructure leverages an existing gaming framework to (i) require understanding of the coordinator for a multi-agent system, (ii) collaborate with human players via instructions, and (iii) enable in-context learning based on few-shot prompting with feedback. Furthermore, we introduce CuisineWorld, a new gaming scenario and its related benchmark that features a multi-agent collaboration efficiency and supervises multiple agents playing the game simultaneously. We have conducted comprehensive evaluations with a new auto-metric collaboration score CoS for assessing the collaboration efficiency. Finally, MindAgent can be deployed in real-world gaming scenarios in a customized VR version of CuisineWorld and adapted in the broader "Minecraft" gaming domain. Our work involving LLMs within our new infrastructure for general-purpose scheduling and coordination can elucidate how such skills may be obtained by learning from large language corpora.
Author context
Most prolific author: 13 submissions (credibility 0.86).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 41% of matchups.
- ▼ lost to Zero-Shot Robotic Manipulation with Pre-Tr… ×4
- ▼ lost to HeaP: Hierarchical Policies for Web Action… ×4
- ▼ lost to The Power of the Senses: Generalizable Man… ×4
- ▼ lost to Tree-Planner: Efficient Close-loop Task Pl… ×4
- ▼ lost to THOUGHT PROPAGATION: AN ANALOGICAL APPROAC… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)