Sep 2026· Journal of Signal Processing Systems· Vol 98· 0 citations· 38 references
TL;DR
This work presents a hardware-aware workflow that combines accurate latency modeling and design space exploration to optimize both neural network architectures and the underlying systolic-array-based accelerator, and demonstrates its approach on ResNet-like networks.
Abstract
When executing common Neural Networks (NNs) on custom AI accelerators, the high performance suggested by advertised Giga or Tera Operations per Second (GOPS/TOPS) is typically not achieved, as low hardware utilization often leads to an effective performance in the single-digit percentage range of the theoretical peak. This discrepancy arises as NNs are typically designed without accounting for the target hardware, leading to inefficient mappings and software optimizations that fail to deliver the expected gains. Addressing this, we present a hardware-aware workflow that combines accurate latency modeling and design space exploration to optimize both neural network architectures and the underlying systolic-array-based accelerator. We develop and validate two high-precision latency models for two different Row-Stationary (RS) dataflows on our target accelerator. Using these models together with a structured search space generation, we generate Pareto-optimal search spaces in terms of achieved GOPS and latency for a given hardware target, and use these for a Bayesian Bayesian Hardware-Aware Neural Architecture Search. We further explore the accelerator design itself in a subsequent hardware DSE stage, varying the PE array dimensions and clock ratio to identify hardware configurations that maximize efficiency and minimize inference latency for each network and dataflow. We demonstrate our approach on ResNet-like networks. On the original hardware, the discovered ResNet-50-like architecture achieves an 85% relative increase in hardware utilization, reduces latency by 21% and parameter count by 18%, and maintains baseline ImageNet accuracy. On the optimized hardware, latency and area are further reduced while efficiency is increased by up to 94%, with up to 20% fewer PEs compared to the baseline configuration. For ResNet-34, similar trends are observed, with latency reductions exceeding 33% and efficiency gains up to 43%. To enable reproducibility, we open-source our complete workflow of latency models, search space generation, NAS and training pipeline.
This review surveys the full landscape of AI hardware accelerators for LLMs, including general-purpose GPUs, custom ASICs, reconfigurable FPGAs, processing-in-memory and near-memory architectures, and emerging neuromorphic and photonic approaches across cloud and edge deployment.
This thesis presents design methodologies that increase the extent to which key design choices are based on exploration and quantitative evaluation of design alternatives, and identifies architectures that consistently outperform the state-of-the-art, achieving considerable improvements in latency, throughput, energy,...
FSGen is proposed, an agile framework for attention-based LLM accelerator generation with an early-stage PPA estimator that supports fused operator dataflows and sparsity with a diverse design space and finds designs with 1.4x better power efficiency or 10x speedup with similar PPA metrics compared to prior work.
A novel open-source framework named OSCAR is proposed, which, given a set of hardware and workload specifications, provides architecture-level power estimation and can also automatically generate Chisel and synthesizable RTL of the custom AI chip.
J. Mok, Qi-Jun Zhang, Di Pang et al.· ACM Transactions on Design A...· 0 citations