NeuroFlex is the first accelerator to assign every output element independently to ANN or SNN execution mode with zero accuracy loss, and extends integer-exact ANN-SNN equivalence from layers to individual output elements, thereby enabling mode switching with no conversion error.
Abstract
Sparse DNN accelerators specialize in ANN or SNN execution, leaving energy or latency on the table when workload characteristics vary within a layer. Hybrid accelerator designs that switch modes at layer or tile granularity suffer from low PE utilization since one core type idles whenever the other is active. NeuroFlex is the first accelerator to assign every output element independently to ANN or SNN execution mode with zero accuracy loss. We extend integer-exact ANN-SNN equivalence from layers to individual output elements, thereby enabling mode switching with no conversion error. An offline cost-guided scheduler scores each element by its marginal energy-delay trade-off and packs work across PEs, achieving 97-99% PE utilization compared to 40-45% for layer-wise hybrids. NeuroFlex reduces EDP by 57-67% over a strong ANN-only baseline and delivers up to 2.5x speedup over a dual-sparse SNN-only baseline. Our cost-guided scheduler improves throughput by 16-19% over random element assignment across vision, language, and transformer workloads.
APEX is presented, a dual-sparsity SNN inference accelerator that integrates the PASC-IF neuron into the LoAS hardware framework, and guarantees mathematical equivalence between the converted SNN and the source ANN, thereby achieving ANN-equivalent accuracy at significantly reduced timesteps.
D. Venkatesh, S. Radhakrishnan, Rajshekhar Rakshit et al.· 0 citations
This work presents a hardware-aware workflow that combines accurate latency modeling and design space exploration to optimize both neural network architectures and the underlying systolic-array-based accelerator, and demonstrates its approach on ResNet-like networks.
Annina Gutermann, Alexey Serdyuk, Foivos Paraskevas et al.· Journal of Signal Processing...· 0 citations
This work proposes a method to distribute the execution of individual layers across accelerators, and demonstrates how to implement such a baseline system using a SoC generator framework, performs an ablation study prototyping different versions on an FPGA, and identifies gaps and limitations by executing a multi-DNN a...
Federico Nicolás Peccia, Avik Bhatnagar, Oliver Bringmann· WiPiEC Journal - Works in Pr...· 0 citations
The 194M-parameter model is implemented on an Alveo U50C FPGA using digital fixed-point arithmetic and on an ARM CPU using sparse integer execution to connect event sparsity to omitted computation and data movement in SymbolicLight V2.
Deep neural networks (DNNs) with billions of parameters power many important applications, but their training is fundamentally constrained by the limited on-chip memory of GPUs. This memory wall forces training to rely on distributed execution or memory offloading, both of which introduce substantial inefficiencies. Ex...
Xiaoyang Sun, Jie Xu, Zheng Wang· IEEE Transactions on Paralle...· 0 citations
Conventional compute-in-memory (CIM) deployment flows usually assume a fixed crossbar geometry, although DNN layers often have different channel counts and matrix shapes. The resulting shape mismatch leaves part of the array capacity unused and can increase the number of split-and-transfer operations during inference....
Kaiwen Deng, Sifan Sun, Hanjie Liu et al.· IEEE Non-Volatile Memory Sys...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.