Dispatching Ambulances using Deep Reinforcement Learning
jon39334@gmail.com, Ole Jakob Mengshoel
OpenReview ground truth
TL;DR — We develop an ambulance-dispatching method based on Proximal Policy Optimization and demonstrate outperformance relative to commonly-used heuristic policies such as dispatching the closest ambulance by Haversine or Euclidean distance.
Abstract
Emergency Medical Service (EMS) plays an essential role in today's society. One EMS component is ambulance dispatch, which impacts the ambulance's response time for a medical incident. Fast response times are essential. The problem of ambulance dispatching differs from a typical Vehicle Routing Problem (VRP) since patients arrive stochastically, making the problem hard to solve. In addition to minimizing response time, EMS providers seek optimal resource utilization and good working conditions for EMS personnel while often experiencing an increase in demand. To meet these requirements, this work develops a Reinforcement learning (RL) method based on Proximal Policy Optimization (PPO) for the ambulance dispatching problem. Varying incident priorities along with more flexible incident queue management are also integrated into our novel method. Our PPO-based method and an EMS simulation model are implemented in Python and combined with Open Street Map (OSM) travel time estimation and simple synthetic incident data generation. Empirical results are presented using both synthetic and real incident data. Results using real incident data from the Oslo University Hospital (OUH) in Norway suggest that our PPO model outperforms heuristic policies such as dispatching the closest ambulance by Haversine or Euclidean distance. We hope that this work inspires future research on RL for ambulance dispatch and ultimately leads to improved decision-support tools for EMS in Norway and elsewhere.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 40% of matchups.
- ▲ beat NL2ProGPT: Taming Large Language Model for… ×8
- ▲ beat Safeguarding Data in Multimodal AI: A Diff… ×6
- ▼ lost to RLP: A reinforcement learning benchmark fo… ×6
- ▼ lost to Interpreting Adaptive Gradient Methods by … ×6
- ▼ lost to Compound Returns Reduce Variance in Reinfo… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)