Skip to content
Open access

Hardware-Aware Early Termination for Low-Latency Spiking Neural Network Inference

2026 · IEEE Access · Vol 14, pp. 148150-148163 · 0 citations · 32 references

Abstract

Spiking Neural Networks (SNNs) provide a natural computation model for neuromorphic hardware, but fixed-timestep inference can execute substantial redundant temporal computation. This work proposes a hardware-aware early termination (ET) framework that determines the stopping time from accumulated output-spike statistics without exporting internal continuous neuron states to the termination controller. A spike-ratio confidence metric is combined with a minimum observation timestep to suppress low-evidence early decisions. To improve the temporal quality of the output spikes used by ET, Time-Weighted Spike Alignment Training (TW-SAT) applies stronger supervision to earlier instantaneous output spikes through a normalized exponentially decaying schedule. Experiments on MNIST, Fashion-MNIST, N-MNIST, and DVS128 Gesture are repeated with three independently trained models. At comparable accuracy, dynamic ET consistently requires fewer executed timesteps than validation-selected fixed cutoffs. On MNIST, ET with TW-SAT reaches <inline-formula> <tex-math notation="LaTeX">$98.53\pm 0.05\%$ </tex-math></inline-formula> accuracy with an average of <inline-formula> <tex-math notation="LaTeX">$2.04\pm 0.01$ </tex-math></inline-formula> timesteps under an algorithm-level <inline-formula> <tex-math notation="LaTeX">$T_{\max }=100$ </tex-math></inline-formula> setting. A complete FPGA implementation on an Artix-7 device preserves timing closure with only 1.68% additional LUTs and 1.86% additional FFs. Under the deployed <inline-formula> <tex-math notation="LaTeX">$T_{\max }=20$ </tex-math></inline-formula> hardware configuration, cycle-accurate simulation shows that ET reduces the average latency from <inline-formula> <tex-math notation="LaTeX">$12.74~\mu $ </tex-math></inline-formula>s to <inline-formula> <tex-math notation="LaTeX">$3.35~\mu $ </tex-math></inline-formula>s and increases effective throughput from 78.5 to 298.5 kFPS, corresponding to a <inline-formula> <tex-math notation="LaTeX">$3.80\times $ </tex-math></inline-formula> improvement. These results show that output-spike-based termination provides a low-overhead control path for reducing SNN inference latency while retaining a spike-observable hardware interface.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.