PapersWithELO
← ICLR 2024 leaderboard

From Images to Connections: Can DQN with GNNs learn the Strategic Game of Hex?

Yannik Keller, Jannis Blüml, Gopika Sudhakaran, Kristian Kersting

reinforcement learningSelf-play Reinforcement LearningGraph Neural NetworksHexLong Range Dependency ProblemsBoard games
11.30100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
13.80100
Mimo
band ≈ ±21 pct pts (from σ = 0.42)
8.60100
DeepSeek
band ≈ ±18 pct pts (from σ = 0.36)

OpenReview ground truth

Rejected

TL;DR — We have studied the use of GNNs instead of CNNs to approximate policy and value functions in the board game Hex, specifically analyzing the impact on the handling of long range dependency problems.

Abstract

The gameplay of strategic board games such as chess, Go and Hex is often characterized by combinatorial, relational structures---capturing distinct interactions and non-local patterns---and not just images. Nonetheless, most common self-play reinforcement learning (RL) approaches simply approximate policy and value functions using convolutional neural networks (CNN). A key feature of CNNs, is their relational inductive biases towards locality and translational invariance. In contrast, graph neural networks (GNN) can encode more complicated and distinct relational structures. Hence, we investigate the crucial question: Can GNNs, with their ability to encode complex connections, replace CNNs in self-play reinforcement learning? To this end, we do a comparison with Hex---an abstract yet strategically rich board game---serving as our experimental platform. Our findings reveal that GNNs excel at dealing with long range dependency situations in game states and are less prone to overfitting, but also showing a reduced proficiency in discerning local patterns. This suggests a potential paradigm shift, signaling the use of game-specific structures to reshape self-play reinforcement learning.

Author context

Most prolific author: 9 submissions (credibility 0.97).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 40 comparisons

Ranked above opponent in 41% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 40)