PapersWithELO
← ICLR 2024 leaderboard

A Hierarchical Reinforcement Learning Based Optimization FrameWork for Large Scale Storage Location Assignment Problem

Weihang Pan, Weixin Xu, YaFei Wang, Yuxiang Zhang

reinforcement learningHierarchical Reinforcement LearningStorage Location Assignment ProblemLarge Scale
22.20100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
22.60100
Mimo
band ≈ ±20 pct pts (from σ = 0.39)
21.30100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

Abstract

The Storage Location Assignment Problem(SLAP) is one of the essential problems within the domain of logistics. The objective is to dynamically allocate optimal storage locations to incoming items, aiming to maximize warehouse space utilization and operational efficiency. Prior research primarily focused on offline scenarios with predetermined goods arrival times. A smaller portion explored real-time allocation using heuristic algorithms based on manual rules and search methods. However, these methods suffer from inadequate solution quality and efficiency, particularly for large-scale problems. To overcome this limitation, we draw inspiration from the partitioned, multi-layered, and modularized layout commonly adopted in most large-scale storage spaces. Building upon this inspiration, we propose a novel hierarchical optimization framework to solve large-scale SLAPs better via reinforcement learning. Specifically, we designed a two-level model: (1) a higher-level model learns to determine which block to choose, and (2) a lower-level model learns to select the final storage location under the constraints of the selected blocks in the upper level. We have designed a policy network based on attention mechanisms for SLAP to achieve better performance. To verify the effectiveness of the proposed framework, we collected a large amount of real historical data from the terminal operating system of Ningbo-Zhoushan Port and built a realistic container terminal simulator. Besides, we conducted extensive offline simulations and online testing using the simulator based on real data and validated the superior performance of our framework compared to existing benchmark methods.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 42% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)