Skip to content

A Flexible Framework for Layer-Parallel CNN Training on FPGA Clusters

Aug 2026 · ACM Transactions on Reconfigurable Technology and Systems · 1 citation · 42 references

TL;DR

A flexible and scalable hardware/software framework for training CNNs on Ethernet-connected FPGA clusters using tightly pipelined layer parallelism, highlighting the potential of FPGA-based, layer-parallel training for separable-convolution-dominated CNNs.

Abstract

We present a flexible and scalable hardware/software framework for training CNNs on Ethernet-connected FPGA clusters using tightly pipelined layer parallelism. Starting from a high-level CNN and cluster description, the system automatically maps layers onto (potentially heterogeneous) devices, generates FPGA-specific bitstreams, and orchestrates fully streaming forward and backward passes while keeping most parameters and gradients in on-chip memory. We demonstrate support for general DAG-style CNNs, including MobileNetV2, MnasNet, and ResNet18, and replace batch normalization with online normalization to enable normalization in this streaming setting while achieving ImageNet validation accuracies comparable to PyTorch baselines with batch normalization for all three networks. A CP-SAT-based planner, driven by implementation-level resource estimates from SpinalHDL, performs resource-aware placement under constraints on DSPs, on-chip memory, DRAM bandwidth, and network bandwidth, and exposes FPGA-specific optimizations such as multipumped Matrix Multiplication engines, fabric-aware memory tiling, and activation recomputation. We evaluate throughput, resource utilization, and energy efficiency on an eight-board Altera Agilex 7 cluster and show that, for both MobileNetV2 and MnasNet, the framework achieves more than \(2\times\) lower energy per frame than Nvidia V100, A100, and H100 GPU baselines, highlighting the potential of FPGA-based, layer-parallel training for separable-convolution-dominated CNNs.

View source

Similar papers

Review Open access Aug 2026

Analysis of Research Progress on Deployment Methods for Deep Learning Models on FPGAs

A systematic review of FPGA-based DL deployment from a cross-layer perspective spanning model, compiler, architecture, runtime, and electronic design automation (EDA) is presented, highlighting that reliable cross-study comparison requires careful consideration of model configuration, precision, execution phase, batch...

Shuo Wang, Lei Chen, Chunsheng Tian et al. · 0 citations
Aug 2026

Hardware-Aware Neural Network Deployment on Multi-Core in-Memory Computing Systems: A Compiler Perspective

Conventional compute-in-memory (CIM) deployment flows usually assume a fixed crossbar geometry, although DNN layers often have different channel counts and matrix shapes. The resulting shape mismatch leaves part of the array capacity unused and can increase the number of split-and-transfer operations during inference....

Kaiwen Deng, Sifan Sun, Hanjie Liu et al. · 0 citations
2026

Automatic Model Compression and Quantized Deployment of Convolutional Neural Networks on Programmable Data Planes

The rapid development of programmable network devices and the widespread adoption of machine learning (ML) in networking have facilitated efficient research into intelligent data planes (IDPs). Offloading ML to programmable data planes (PDPs) enables quick analysis and responses to network traffic dynamics, and efficie...

Xiaoquan Zhang, Mai Zhang, L. Cui et al. · 0 citations
Conference Aug 2026

Designing and Building an FPGA Accelerator That Uses Less Energy for DNN Inference

Deep Neural Networks (DNNs) are critical to modern AI applications, yet their deployment on standard CPUs and GPUs is constrained by high power consumption and computational latency, particularly in resource-constrained edge environments. To address these limitations, this paper presents the design and implementation o...

P. V. G. K. Rao, Dudekula Raziya · 0 citations
Conference Aug 2026

FPGA-Based Hardware Accelerator for U-Net: A Resource-Efficient, Pipelined Micro-Architecture for Real-Time Image Segmentation

The increasing adoption of modern embedded platforms, edge devices, and AI driven systems has led to higher computational demands. To facilitate that, there should be hardware acceleration techniques capable of delivering higher throughput with minimal latency. Most of the traditional hardware accelerator architectures...

K. O. Y. N. Karunanayake, H. D. I. J. A. Deshapriya, A. T. Saiamirthan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.