From Images to Connections: Can DQN with GNNs learn the Strategic Game of Hex?
Yannik Keller, Jannis Blüml, Gopika Sudhakaran, Kristian Kersting
OpenReview ground truth
TL;DR — We have studied the use of GNNs instead of CNNs to approximate policy and value functions in the board game Hex, specifically analyzing the impact on the handling of long range dependency problems.
Abstract
The gameplay of strategic board games such as chess, Go and Hex is often characterized by combinatorial, relational structures---capturing distinct interactions and non-local patterns---and not just images. Nonetheless, most common self-play reinforcement learning (RL) approaches simply approximate policy and value functions using convolutional neural networks (CNN). A key feature of CNNs, is their relational inductive biases towards locality and translational invariance. In contrast, graph neural networks (GNN) can encode more complicated and distinct relational structures. Hence, we investigate the crucial question: Can GNNs, with their ability to encode complex connections, replace CNNs in self-play reinforcement learning? To this end, we do a comparison with Hex---an abstract yet strategically rich board game---serving as our experimental platform. Our findings reveal that GNNs excel at dealing with long range dependency situations in game states and are less prone to overfitting, but also showing a reduced proficiency in discerning local patterns. This suggests a potential paradigm shift, signaling the use of game-specific structures to reshape self-play reinforcement learning.
Author context
Most prolific author: 9 submissions (credibility 0.97).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 40 comparisons
Ranked above opponent in 41% of matchups.
- ▲ beat State Representation Learning Using an Unb… ×6
- ▼ lost to Cosine Similarity Knowledge Distillation f… ×6
- ▲ beat Video Generation Beyond a Single Clip ×6
- ▼ lost to Detecting Influence Structures in Multi-Ag… ×4
- ▼ lost to Robust NAS under adversarial training: ben… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 40)