A Hardware-Efficient SRAM-Based CNN Accelerator for Edge Image Classification
Abstract
CNN inference on edge hardware is constrained by memory bandwidth, redundant logic, and the area overhead of fully parallel multiply-accumulate (MAC) units. This paper presents a hardware-efficient CNN accelerator with an SRAM-based architecture optimised for classifying 28×28 grayscale images. Input images are ingested through line buffers and a sliding-window extractor that supply reusable 3×3 patches to an 8-lane convolution engine. Each lane implements partial-parallel 4-MAC processing, accumulation, ReLU activation, requantisation, and max-pooling, achieving output-channel parallelism while keeping hardware resources under control. To further reduce area, the second convolutional stage reuses the same 8-lane engine in two passes to realise 16 filters, after which a flattening layer, a fully connected layer, and argmax-based classification complete inference. Post-synthesis results show that the design meets a 10ns timing constraint with no setup violations, a total area of 2,219,599.07units, and 42,825 leaf cells. Total power consumption is 127.35mW, with on-chip memory accounting for 98.78% of that figure, indicating that future optimisation should focus on SRAM access scheduling to reduce energy consumption in alignment with sustainable design principles.