Temporal Parallelization for GPU Acceleration of Spiking Neural Networks
Jiachun Li, Yanchen Li, Kebin Sun, Ran Cheng
OpenReview ground truth
Abstract
Inspired by neurobiological structures, Spiking Neural Networks (SNNs) are heralded as a significant advancement in deep learning, given their potential for superior computational efficiency. However, this potential often remains untapped on contemporary hardware platforms. Specifically, when deployed on standard GPUs, SNNs tend to require extended computation times, placing them at a disadvantage compared to traditional Artificial Neural Networks (ANNs). Such inefficiencies have somehow diminished enthusiasm for SNN research and presented the tangible challenge to achieving scalability. To address such a challenge, this study introduces a temporal parallelization method specifically tailored for accelerating the propagation dynamics of SNNs on GPUs. Furthermore, we furnish two distinct implementations\footnote{The source code will be made publicly available.} based on the CUDA and JAX frameworks respectively, ensuring adaptability across both single and multi-GPU setups. When benchmarked against several established SNN implementations, the empirical analysis confirmed the efficacy of our proposed method. Notably, with the Leaky Integrate-and-Fire model as a test case, the CUDA-based implementation achieved $5\times$ to $40\times$ acceleration on the A100 GPU.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 39% of matchups.
- ▼ lost to MoteS: Memory Optimization via Fine-graine… ×6
- ▲ beat Automated Search-Space Generation Neural A… ×6
- ▼ lost to Adversarial Defense using Targeted Manifol… ×6
- ▼ lost to Lion Secretly Solves a Constrained Optimiz… ×4
- ▼ lost to FIITED: Fine-grained embedding dimension o… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)