PapersWithELO
← ICLR 2024 leaderboard

MindAgent: Emergent Gaming Interaction

Ran Gong, Qiuyuan Huang, Xiaojian Ma, Hoi Vo, Zane Durante, Yusuke Noda, Zilong Zheng, Demetri Terzopoulos, Li Fei-Fei, Jianfeng Gao

robotics & planningLarge Language ModelsDecision MakingMulti-Agent SystemsGaming Interaction
12.20100
Fused
band ≈ ±15 pct pts (from σ = 0.29)
12.10100
Mimo
band ≈ ±23 pct pts (from σ = 0.45)
13.60100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.37)

OpenReview ground truth

Rejected

Abstract

Large Language Models (LLMs) can perform complex scheduling in a multi-agent system and can coordinate agents to complete sophisticated tasks that require extensive collaboration. However, despite the introduction of numerous gaming frameworks, the community lacks adequate benchmarks that support the implementation of a general multi-agent infrastructure encompassing collaboration between LLMs and human-NPCs. We propose a novel infrastructure--- MindAgent---for evaluating planning and coordination-emergent capabilities in the context of gaming interaction. In particular, our infrastructure leverages an existing gaming framework to (i) require understanding of the coordinator for a multi-agent system, (ii) collaborate with human players via instructions, and (iii) enable in-context learning based on few-shot prompting with feedback. Furthermore, we introduce CuisineWorld, a new gaming scenario and its related benchmark that features a multi-agent collaboration efficiency and supervises multiple agents playing the game simultaneously. We have conducted comprehensive evaluations with a new auto-metric collaboration score CoS for assessing the collaboration efficiency. Finally, MindAgent can be deployed in real-world gaming scenarios in a customized VR version of CuisineWorld and adapted in the broader "Minecraft" gaming domain. Our work involving LLMs within our new infrastructure for general-purpose scheduling and coordination can elucidate how such skills may be obtained by learning from large language corpora.

Author context

Most prolific author: 13 submissions (credibility 0.86).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 41% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)