Low-Power Binary Neural Network Accelerator with Scalable Processing and Bit-Stream Input
Abstract
This paper presents the design, simulation, and ASIC-focused implementation of a Binary Convolutional Neural Network (BCNN) tailored for energy-efficient edge AI applications. The proposed architecture utilizes binarized weights and activations, substituting standard multipliers with XNOR-popcount operations to reduce hardware complexity and energy usage. The system utilizes modular Verilog HDL, combining convolution and thresholding units in a pipelined architecture to achieve high processing efficiency. The design features parallel processing lanes and clock-gating methods to improve throughput and lower dynamic power usage. Functional verification validates the accurate of binary inference, while post-synthesis findings from Cadence Genus show decreased area and power relative to traditional convolution neural network implementations. The proposed BCNN architecture shows the feasibility of binary neural networks for real-time, low-energy edge intelligence and offers a scalable and reusable framework for future ASIC-driven deep learning accelerators