Skip to content
Book Open access

Algorithm--Hardware Co-Design of Spiking Transformers for In-Memory Neuromorphic Edge Vision

Sep 2026 · Proceedings of the International Conference on Parallel Processing · pp. 436-445 · 0 citations · 14 references

Abstract

Reliable, ultra-low-power neuromorphic vision on edge devices requires attention mechanisms that combine accuracy with parallel, memory-efficient execution. Current Spiking Transformers suit neuromorphic data but retain the \(\mathcal {O}(N^2D)\) complexity of standard self-attention and weak temporal modeling, limiting their efficiency on in-memory edge hardware. We propose the Temporal Hadamard Transformer (THT), a spiking model co-designed with hardware for in-memory edge vision. Its Temporal Hadamard Attention (THA) combines a lightweight Temporal Processing Unit (TPU) for short-term temporal fusion with a binary Mask&Add operator. By replacing Query–Key–Value computation with accumulate-and-fire logic that maps efficiently onto memristor crossbars, THA reduces attention complexity to \(\mathcal {O}(ND)\) and enables highly parallel in-memory execution. On CIFAR10-DVS, THT achieves a peak Accuracy-to-Energy ratio of 157.06. THT consumes 0.51–2.19 mJ per inference. Its lowest-energy configuration uses up to 16 × lower energy than the compared models, while the largest configuration reaches 81.8% accuracy. We further design a detailed hardware deployment scheme and evaluate its robustness through non-ideal memristor crossbar simulations. Accounting for conductance drift and IR drop, the simulated deployment maintains 81.1% accuracy. Thus, THT integrates sparse algorithms with parallel hardware execution, providing a practical path toward real-time edge systems under realistic resource constraints and physical device non-idealities.

Read PDF

Similar papers

Open access Oct 2026

Beyond binary SNNs: a hardware-algorithm co-design and architectural evaluation for graded spike networks

While Spiking Neural Networks offer a promising path toward energy-efficient edge intelligence, conventional Binary representations often suffer from high inference latency and discretization errors. Multi-level (graded) spike models address these limitations by increasing information density per pulse, yet they typi...

Wen-Fei Song, Andrea Castagnetti, Pierre-Emmanuel Novac et al. · 0 citations
#machine learning Preprint Sep 2026

MorphAtt: A Neuromorphic Accelerator for Efficient Multi-Head Attention Processing in Spiking Vision Transformers

Spiking Vision Transformers (SViTs) are developed as an energy-efficient alternative to conventional ViTs for computer vision tasks at the edge. However, huge parameter counts and complex multi-head self-attention (MHSA) operations make it challenging to achieve high energy efficiency in SViT inference, especially in t...

Rachmad Vidya Wicaksana Putra, Amirhesam Jafari Rad, M. Shafique · 0 citations
Review Open access Sep 2026

Spiking Neural Network Chips for Low-Power Real-Time Edge Intelligence

Edge devices require short response time and low active and idle power, yet conventional neural-network hardware repeatedly moves weights and activations between memory and processing units. Spiking neural network (SNN) chips address this problem with event-driven communication, persistent neuron state, and local synap...

Le-Yang Qin · 0 citations
#edge computing Review Open access Sep 2026

GPU and RISC-V acceleration for neuromorphic computing based on spiking neural networks: taxonomy, comparison, and open challenges

This survey reviews 71 studies published between 2015 and 2025 and organizes them into a taxonomy that classifies GPU works into simulation frameworks, training-acceleration techniques, large-scale and multi-GPU simulation, and application deployments, and RISC-V works into instruction-set extensions and tightly-couple...

Edris Zaman Farsa, Amirhossein Ilkhani, Marc Reichenbach et al. · 0 citations
Open access Aug 2026

A hardware-implemented adaptive pruning neural network

With the rapid growth of artificial intelligence (AI) applications, there is an urgent demand for edge computing hardware with low latency and high energy efficiency. The human brain, a highly parallel and sparsely connected architecture, can efficiently process complex tasks at exceptionally low power consumption, a...

Yusen Tian, Penghao Chen, Ziyu Ming et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.