PapersWithELO
← ICLR 2024 leaderboard

Learning Team-Level Information Integration in Multi-Agent Communication

Xiangrui Meng, Ying Tan

reinforcement learningMulti-Agent CommunicationMulti-Agent Reinforcement LearningTeam-Level CommunicationDeep Reinforcement Learning
14.70100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
15.90100
Mimo
band ≈ ±19 pct pts (from σ = 0.38)
10.70100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

TL;DR — Double Channel Communication Model

Abstract

In human cooperation, both individual knowledge and group consensus play important roles in accomplishing tasks. However, existing multi-agent reinforcement learning (MARL) communication methods commonly focus on individual-level communication, which lacks the necessary global information for well-grounded decision-making. Meanwhile, individual-level communication is often infeasible when the communication bandwidth is limited. To tackle these problems, we propose a group-level information integration model called Double Channel Communication Network (DC2Net). DC2Net highlights the significance of independent group feature learning by separating individual and group feature learning into two independent channels. In this model, agents no longer communicate with each other in a peer-to-peer paradigm; instead, all interactions are carried out in the group channel. By combining individual and global features, decisions are made collaboratively. We conduct experiments on several multi-agent cooperative environments and the results show that the DC2Net not only outperforms state-of-the-art MARL communication models but also reduces the communication costs. Furthermore, the two independent channels enable adaptive balancing of individual and group feature learning based on task requirements.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)