Skip to content
Conference

FPGA-Based Hardware Accelerator for U-Net: A Resource-Efficient, Pipelined Micro-Architecture for Real-Time Image Segmentation

Aug 2026 · Moratuwa Engineering Research Conference · pp. 628-633 · 0 citations · 10 references

Abstract

The increasing adoption of modern embedded platforms, edge devices, and AI driven systems has led to higher computational demands. To facilitate that, there should be hardware acceleration techniques capable of delivering higher throughput with minimal latency. Most of the traditional hardware accelerator architectures rely on either High Level Synthesis Tools or utilizing already available Deep Learning Units, which often cost a huge portion of the hardware resources available on a Field Programmable Gate Array. This paper presents a resource efficient pipelined micro architecture optimized for U Net Convolutional Neural Network Architecture for image segmentation, targeting improved resource efficiency while allowing data level parallelism. The proposed design addresses resource efficiency via designing hardware units on Register Transfer Level while using smaller data sizes. Mathematical operations are divided into multiple pipeline stages, which enable the concurrent processing of data streams. This work provides a flexible technique for designing high performance hardware accelerators for real time image processing workloads.

View source

Similar papers

Aug 2026

A novel design of high throughput power efficient multiply accumulate unit

The proposed Double MAC unit with dynamic precision scaling has showed twofold improvement in throughput and 15% improvement in power consumption, and proves advantageous in convolution layers, where greater precision is required for final classification and smaller precision in initial stage.

M. Jayasanthi, R. Kalaivani, K. P. Sampoornam et al. · 0 citations
Conference Aug 2026

Designing and Building an FPGA Accelerator That Uses Less Energy for DNN Inference

Deep Neural Networks (DNNs) are critical to modern AI applications, yet their deployment on standard CPUs and GPUs is constrained by high power consumption and computational latency, particularly in resource-constrained edge environments. To address these limitations, this paper presents the design and implementation o...

P. V. G. K. Rao, Dudekula Raziya · 0 citations
Conference Aug 2026

A Hardware-Efficient SRAM-Based CNN Accelerator for Edge Image Classification

CNN inference on edge hardware is constrained by memory bandwidth, redundant logic, and the area overhead of fully parallel multiply-accumulate (MAC) units. This paper presents a hardware-efficient CNN accelerator with an SRAM-based architecture optimised for classifying 28×28 grayscale images. Input images are ingeste...

Avinash Krishna Pk, R. S · 0 citations
Conference Aug 2026

An Efficient HLS-Based Hardware Accelerator with Resource Optimization for Transformer Models

Deploying Transformer models on FPGA and System-on-Chip (SoC) platforms remains challenging due to their substantial computational complexity, large memory footprint, and high hardware resource requirements, particularly in multi-head attention and stacked encoder-decoder layers. This paper proposes a hardware-efficien...

Xuan Thao Tran, Thi Diem Tran · 0 citations
2026

Modern Implementation of Image Processing Algorithms on FPGA

Field-programmable gate arrays (FPGAS) have emerged as a powerful platform for real-time image processing due to their inherent parallelism and configurability. This paper presents an optimized hardware implementation of fundamental image processing algorithms including Sobel edge detection, Thresholding contrast stret...

Pramod Moud, P. Sharma · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.