Conventional processing-in-memory (PIM) architectures suffer from limited efficiency due to transistor-intensive adder trees and analog-to-digital converter (ADC) overhead. This work presents BRAIN-AD, a bit-serial, ReRAM-based digital PIM macro for energy-efficient autonomous driving assistance systems (ADAS). The proposed design integrates a compact 1-bit multiply-accumulate (MAC) unit combining a 3T1R Re-XNOR non-volatile bit-cell for in-memory multiplication with an area-efficient 10 T pass-transistor full adder for sequential accumulation. A 16 Kb (128 × 128) macro employs sparsity-aware power gating and supports scalable fixed-point computation from 1 to 16 bits via bit-serial execution. Post-layout simulations in 65 nm CMOS achieve peak throughput of 0.72 TOPS and 112 TOPS/W energy efficiency, providing approximately 1.8× higher throughput and 1.95× higher energy efficiency than state-of-the-art digital PIM designs. System-level evaluation using a quantised INT4 NVIDIA PilotNet model shows less than 2.5% accuracy degradation relative to the FP32 baseline. These results establish BRAIN-AD as a robust, scalable, and practical digital PIM solution for resource-constrained ADAS workloads. This work highlights digital ReRAM-based bit-serial PIM as a scalable and robust alternative to analog CIM for safety-critical edge-AI applications.
Ankit Kumar Tenwar, Mukul Lokhande, A. Teman et al.· IEEE transactions on nanotec...· 0 citations
The growing demand for efficient deep-learning inference on edge platforms requires hardware that is both energy-efficient and practically implementable. This work presents a 16-Kb all-digital static random-access memory (SRAM)-based compute-in-memory (CIM) macro for low-bit CNN inference, featuring a hierarchical adder-tree-based accumulation architecture. The design integrates a nor-enabled SRAM compute cell, column-wise rearrangement network, sparsity-aware compression, and multistage hierarchical accumulation within a 64-bank $64 \,\, \times \,\, 4$ architecture, enabling scalable bit-serial processing and utilization-aware mapping. Implemented in 65-nm CMOS, the macro achieves 8.19 TOPS effective throughput at 1.0 V and a peak energy efficiency of 586 TOPS/W at 0.9 V under practical operating conditions. Hardware-compatible CNN mapping is demonstrated using LeNet-5, VGG-8, and ResNet-8. The design achieves 98.1% and 72.3% accuracy on MNIST and CIFAR-10, respectively, with 1-bit activations and 4-bit weights, while 4-bit configurations on deeper networks show only 3%–4% degradation from FP32 baselines. CNN inference is evaluated using a hardware-compatible post-training quantization (PTQ) flow without retraining. These results demonstrate that the proposed SRAM-CIM architecture provides an efficient and scalable accumulation solution with a practical tradeoff among throughput, energy efficiency, and implementability for edge-oriented deep neural network (DNN) inference.
Vikash Vishwakarma, Gopal R. Raut, Amit Mittal et al.· IEEE Transactions on Very La...· 0 citations