PapersWithELO
← ICLR 2024 leaderboard

Learning Multiple Coordinated Agents under Directed Acyclic Graph Constraints

Jaeyeon Jang, Diego Klabjan, Han Liu, Nital S Patel, Xiuqi Li, Balakrishnan Ananthanarayanan, Husam Dauod, Tzung-Han Juang

reinforcement learningMulti-agent reinforcement learningdirected acyclic graphsynthetic rewardreward shaping
42.90100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
41.70100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
49.00100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

Abstract

This paper proposes a novel multi-agent reinforcement learning (MARL) method to learn multiple coordinated agents under directed acyclic graph (DAG) constraints. Unlike existing MARL approaches, our method explicitly exploits the DAG structure between agents to achieve more effective learning performance. Theoretically, we propose a novel surrogate value function based on a MARL model with synthetic rewards (MARLM-SR) and prove that it serves as a lower bound of the optimal value function. Computationally, we propose a practical training algorithm that exploits new notion of leader agent and reward generator and distributor agent to guide the decomposed follower agents to better explore the parameter space in environments with DAG constraints. Empirically, we exploit four DAG environments including a real-world scheduling for one of Intel’s high volume packaging and test factory to benchmark our methods and show it outperforms the other non-DAG approaches.

Author context

Most prolific author: 4 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 52% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)