Skip to content
Preprint

NeuroFlex: Lossless Element-Level ANN-SNN Co-Execution for Efficient Sparse Inference

Nov 2025 · 0 citations · 27 references
Computer Science

TL;DR

NeuroFlex is the first accelerator to assign every output element independently to ANN or SNN execution mode with zero accuracy loss, and extends integer-exact ANN-SNN equivalence from layers to individual output elements, thereby enabling mode switching with no conversion error.

Abstract

Sparse DNN accelerators specialize in ANN or SNN execution, leaving energy or latency on the table when workload characteristics vary within a layer. Hybrid accelerator designs that switch modes at layer or tile granularity suffer from low PE utilization since one core type idles whenever the other is active. NeuroFlex is the first accelerator to assign every output element independently to ANN or SNN execution mode with zero accuracy loss. We extend integer-exact ANN-SNN equivalence from layers to individual output elements, thereby enabling mode switching with no conversion error. An offline cost-guided scheduler scores each element by its marginal energy-delay trade-off and packs work across PEs, achieving 97-99% PE utilization compared to 40-45% for layer-wise hybrids. NeuroFlex reduces EDP by 57-67% over a strong ANN-only baseline and delivers up to 2.5x speedup over a dual-sparse SNN-only baseline. Our cost-guided scheduler improves throughput by 16-19% over random element assignment across vision, language, and transformer workloads.

View source

Similar papers

Preprint Aug 2026

APEX: A Dual-Sparsity Accelerator for Precise and Efficient SNN Inference

APEX is presented, a dual-sparsity SNN inference accelerator that integrates the PASC-IF neuron into the LoAS hardware framework, and guarantees mathematical equivalence between the converted SNN and the source ANN, thereby achieving ANN-equivalent accuracy at significantly reduced timesteps.

D. Venkatesh, S. Radhakrishnan, Rajshekhar Rakshit et al. · 0 citations
Open access Sep 2026

Joint Exploration of Neural Networks and Systolic Hardware for Improved AI Accelerator Performance

This work presents a hardware-aware workflow that combines accurate latency modeling and design space exploration to optimize both neural network architectures and the underlying systolic-array-based accelerator, and demonstrates its approach on ResNet-like networks.

Annina Gutermann, Alexey Serdyuk, Foivos Paraskevas et al. · 0 citations
Open access Aug 2026

Autotuned Distribution of Multi-DNN Workloads on Multi-Accelerator SoCs

This work proposes a method to distribute the execution of individual layers across accelerators, and demonstrates how to implement such a baseline system using a SoC generator framework, performs an ablation study prototyping different versions on an FPGA, and identifies gaps and limitations by executing a multi-DNN a...

Federico Nicolás Peccia, Avik Bhatnagar, Oliver Bringmann · 0 citations
Oct 2026

SynergyScale: Optimizing Offloading and Task Partitioning for Efficient Model Training

Deep neural networks (DNNs) with billions of parameters power many important applications, but their training is fundamentally constrained by the limited on-chip memory of GPUs. This memory wall forces training to rely on distributed execution or memory offloading, both of which introduce substantial inefficiencies. Ex...

Xiaoyang Sun, Jie Xu, Zheng Wang · 0 citations
Aug 2026

Hardware-Aware Neural Network Deployment on Multi-Core in-Memory Computing Systems: A Compiler Perspective

Conventional compute-in-memory (CIM) deployment flows usually assume a fixed crossbar geometry, although DNN layers often have different channel counts and matrix shapes. The resulting shape mismatch leaves part of the array capacity unused and can increase the number of split-and-transfer operations during inference....

Kaiwen Deng, Sifan Sun, Hanjie Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.