Beyond binary SNNs: a hardware-algorithm co-design and architectural evaluation for graded spike networks
Abstract
While Spiking Neural Networks offer a promising path toward energy-efficient edge intelligence, conventional Binary representations often suffer from high inference latency and discretization errors. Multi-level (graded) spike models address these limitations by increasing information density per pulse, yet they typically necessitate power-hungry Multiply-Accumulate units that nullify the energy advantages of neuromorphic hardware. In this work, we propose a hardware-algorithm co-design framework that leverages multi-level Integrate-and-Fire neurons coupled with a specialized multiplier-less architecture. By decomposing multi-level interactions into a cascaded shift-and-add logic, our hardware preserves the event-driven sparsity of SNNs while supporting high-precision activations. We evaluate our proposed accelerator on the Xilinx Zynq MPSoC ZCU102 platform using the Google Speech Commands, CIFAR-10, and DVS-Gesture datasets. Experimental results demonstrate that the graded spike paradigm achieves up to a 6.7× reduction in total spike activity through temporal information compression. Compared with the binary baseline benchmark on the DVS Gesture dataset, the proposed architecture reduces memory accesses by 68.59%, improves steady-state throughput from 33.08 FPS to 251 FPS, and increases energy efficiency from 9.01 to 54.96 GOPS/W. Furthermore, by minimizing pipeline stalls and workload imbalance, our design achieves an Efficiency-per-Processing-Element of up to 3.05 GOPS/W/PE, thereby providing a competitive PE-level energy efficiency compared with representative SNN accelerators. This research provides a scalable pathway for deploying high-performance, ”always-on” neuromorphic systems in resource-constrained environments.