A novel simulated annealing based mapping for hybrid NoC-enabled DNN accelerators
Abstract
With the continuous development of very large-scale integration (VLSI), the number of cores integrated on a single chip has reached hundreds. Network-on-Chip (NoC), featuring a highly scalable and high-bandwidth communication architecture, has been widely applied in Chip Multiprocessor Systems (CMP). NoC-based deep neural network (DNN) accelerators leverage their high-parallelism characteristics to efficiently execute various convolutional computations of neural networks in parallel. However, when handling large-scale neural networks, the traditional mesh topology faces challenges such as excessive packet hops and transmission congestion. To address these issues, this paper proposes a reconfigurable hybrid ring-shaped architecture (Rhr-NoC) to adapt to the large-scale data transmission patterns within accelerators. Additionally, we mathematically model the mapping task and resource allocation in NoCs and design a simulated annealing algorithm tailored to the hybrid ring-shaped architecture, providing a more rational scheme for mapping neural networks onto NoC platforms. Detailed experimental results show that our scheme outperforms baseline approaches. Specifically, compared to traditional mesh topologies and sequential mapping, our proposed co-design reduces the average packet delay by up to 15.6%, decreases the classification latency by 10.46% on average, and achieves a 16.99% reduction in total dynamic power consumption, significantly improving the system transmission efficiency.