← ICLR 2024 leaderboard

Enhancing Sample Efficiency in Black-box Combinatorial Optimization via Symmetric Replay Training

Hyeonah Kim, Minsu Kim, Sungsoo Ahn, Jinkyoo Park

reinforcement learningBlack-box combinatorial optimizationsample efficiencysymmetriesdrug discoveryhardware designdeep reinforcement learningimitation learning
28.80100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
35.80100
Mimo
band ≈ ±19 pct pts (from σ = 0.37)
23.10100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.43)

OpenReview ground truth

Rejected

TL;DR — This paper proposes a general approach to improve the sample efficiency of DRL for black-box combinatorial optimization by exploiting symmetric transformations.

Abstract

Black-box combinatorial optimization (black-box CO) is frequently encountered in various industrial fields, such as drug discovery or hardware design. Despite its widespread relevance, solving black-box CO problems is highly challenging due to the vast combinatorial solution space and resource-intensive nature of black-box function evaluations. These inherent complexities induce significant constraints on the efficacy of existing deep reinforcement learning (DRL) methods when applied to practical problem settings. For efficient exploration with the limited availability of function evaluations, this paper introduces a new generic method to enhance sample efficiency. We propose symmetric replay training that leverages the high-reward samples and their under-explored regions in the symmetric space. In replay training, the policy is trained to imitate the symmetric trajectories of these high-rewarded samples. The proposed method is beneficial for the exploration of highly rewarded regions without the necessity for additional online interactions - free. The experimental results show that our method consistently improves the sample efficiency of various DRL methods on real-world tasks, including molecular optimization and hardware design. Our source code is available at https://anonymous.4open.science/r/sym_replay.

Author context

Most prolific author: 10 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 38 comparisons

Ranked above opponent in 46% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)