Aug 2026· Moratuwa Engineering Research Conference· pp. 628-633· 0 citations· 10 references
Abstract
The increasing adoption of modern embedded platforms, edge devices, and AI driven systems has led to higher computational demands. To facilitate that, there should be hardware acceleration techniques capable of delivering higher throughput with minimal latency. Most of the traditional hardware accelerator architectures rely on either High Level Synthesis Tools or utilizing already available Deep Learning Units, which often cost a huge portion of the hardware resources available on a Field Programmable Gate Array. This paper presents a resource efficient pipelined micro architecture optimized for U Net Convolutional Neural Network Architecture for image segmentation, targeting improved resource efficiency while allowing data level parallelism. The proposed design addresses resource efficiency via designing hardware units on Register Transfer Level while using smaller data sizes. Mathematical operations are divided into multiple pipeline stages, which enable the concurrent processing of data streams. This work provides a flexible technique for designing high performance hardware accelerators for real time image processing workloads.
The proposed Double MAC unit with dynamic precision scaling has showed twofold improvement in throughput and 15% improvement in power consumption, and proves advantageous in convolution layers, where greater precision is required for final classification and smaller precision in initial stage.
M. Jayasanthi, R. Kalaivani, K. P. Sampoornam et al.· Analog Integrated Circuits a...· 0 citations
Deep Neural Networks (DNNs) are critical to modern AI applications, yet their deployment on standard CPUs and GPUs is constrained by high power consumption and computational latency, particularly in resource-constrained edge environments. To address these limitations, this paper presents the design and implementation o...
P. V. G. K. Rao, Dudekula Raziya· 2026 International Conferenc...· 0 citations
CNN inference on edge hardware is constrained by memory bandwidth, redundant logic, and the area overhead of fully parallel multiply-accumulate (MAC) units. This paper presents a hardware-efficient CNN accelerator with an SRAM-based architecture optimised for classifying 28×28 grayscale images. Input images are ingeste...
Avinash Krishna Pk, R. S· International Conference Inn...· 0 citations
Deploying Transformer models on FPGA and System-on-Chip (SoC) platforms remains challenging due to their substantial computational complexity, large memory footprint, and high hardware resource requirements, particularly in multi-head attention and stacked encoder-decoder layers. This paper proposes a hardware-efficien...
Xuan Thao Tran, Thi Diem Tran· International Conference on...· 0 citations
Field-programmable gate arrays (FPGAS) have emerged as a powerful platform for real-time image processing due to their inherent parallelism and configurability. This paper presents an optimized hardware implementation of fundamental image processing algorithms including Sobel edge detection, Thresholding contrast stret...
Pramod Moud, P. Sharma· International Journal of Lat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.