PapersWithELO
← ICLR 2024 leaderboard

Towards Precise Prediction Uncertainty in GNNs: Refining GNNs with Topology-grouping Strategy

Hyunjin Seo, Kyusung Seo, Joonhyung Park, Eunho Yang

fairness, safety & privacyGraph Neural NetworksPost-hoc Calibration
19.80100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
23.30100
Mimo
band ≈ ±19 pct pts (from σ = 0.39)
16.70100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

Abstract

The calibration of model predictions has recently gained increasing attention in the domain of graph neural networks (GNNs), with a particular emphasis on the underconfidence exhibited by these networks. Among the critical factors identified to be associated with GNN calibration, the concept of neighborhood prediction similarity has been recognized as a pivotal component. Building upon this insight, modern GNN calibration techniques adapt GNNs by smoothing the confidence of individual nodes with those of adjacent nodes. However, these approaches often engage in superficial learning across varying affinity levels, thereby failing to effectively accommodate diverse local topologies. Through an in-depth analysis, we unveil that calibrated logits from preceding research significantly contradict their foundational assumption of nearby affinity, necessitating a re-evaluation of the existing GNN-founded calibration strategies. To address this, we introduce Simi-Mailbox, which categorizes nodes based on both neighborhood representational similarity and their own confidence, irrespective of proximity or connectivity. Our method effectively mitigates miscalibration for nodes exhibiting analogous similarity levels by adjusting their predictions with group-specific temperatures. This encourages a more sophisticated calibration, where each group-wise temperature is tailored to address affiliated nodes with similar topology. Extensive experiments demonstrate the effectiveness of Simi-Mailbox across diverse datasets on different GNN architectures.

Author context

Most prolific author: 8 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 38% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)