Skip to content

FAME: An FPGA-Based Platform for Approximate Multipliers Evaluation with Pattern-Guided DNN Retraining

Sep 2026 · 0 citations · 44 references
Computer Science

TL;DR

The platform exploits the reconfigurable logic of Field-Programmable Gate Arrays to implement approximate multipliers directly in hardware, eliminating the need for LUT-based emulation on CPU/GPU platforms and thereby enabling efficient DNN inference while significantly reducing evaluation time on large datasets.

Abstract

Approximate multipliers can reduce hardware area and energy consumption in Deep Neural Network (DNN) inference; however, they introduce computational errors. Assessing the accuracy of numerous approximate multiplier designs across diverse DNN models and large-scale datasets remains challenging due to prohibitive evaluation times. This overhead primarily stems from the slow emulation of approximate multiplier behavior using look-up tables (LUTs) on CPU and GPU platforms. Moreover, the resulting accuracy degradation must be carefully quantified and, if necessary, mitigated (e.g., through retraining), further increasing the overall evaluation cost. To address these challenges, we propose FAME, an FPGA-based platform for evaluating approximate multipliers. The platform exploits the reconfigurable logic of Field-Programmable Gate Arrays (FPGAs) to implement approximate multipliers directly in hardware, eliminating the need for LUT-based emulation on CPU/GPU platforms and thereby enabling efficient DNN inference while significantly reducing evaluation time on large datasets. Furthermore, we introduce a pattern-guided DNN retraining technique to mitigate accuracy degradation induced by approximate multipliers. Specifically, retraining is guided by multiplier-specific patterns to effectively recover potential accuracy losses. We evaluate FAME using two DNN models, ResNet-18 and MobileNetV2, on the ImageNet dataset across 27 approximate multipliers. During inference, our approach achieves up to a 3.47x speedup in approximate multiplier evaluation compared to prior LUT-based emulation methods. Furthermore, the proposed retraining technique improves accuracy by up to 65.5% over existing retraining approaches for the evaluated multipliers. The code is publicly available at: https://github.com/gicLAB/FAME

View source

Similar papers

Conference Aug 2026

Designing and Building an FPGA Accelerator That Uses Less Energy for DNN Inference

Deep Neural Networks (DNNs) are critical to modern AI applications, yet their deployment on standard CPUs and GPUs is constrained by high power consumption and computational latency, particularly in resource-constrained edge environments. To address these limitations, this paper presents the design and implementation o...

P. V. G. K. Rao, Dudekula Raziya · 0 citations
Review Aug 2026

Model Compression and Hardware-Aware Acceleration for Deep Learning on FPGAs: A Co-Design Taxonomy and Comparative Analysis

This survey restricted the scope to FPGAs alone and organized 25 compression-hardware co-design case studies into a five-category taxonomy defined by which FPGA resources each strategy primarily reshapes, formalizing six open challenges: toolchain fragmentation, accuracy--efficiency characterization, automated mixed-pr...

Peter Forcha, H. Kajekusumadhar, Mbua Peter et al. · 0 citations
Conference Aug 2026

An Efficient HLS-Based Hardware Accelerator with Resource Optimization for Transformer Models

Deploying Transformer models on FPGA and System-on-Chip (SoC) platforms remains challenging due to their substantial computational complexity, large memory footprint, and high hardware resource requirements, particularly in multi-head attention and stacked encoder-decoder layers. This paper proposes a hardware-efficien...

Xuan Thao Tran, Thi Diem Tran · 0 citations
Review Open access Aug 2026

Analysis of Research Progress on Deployment Methods for Deep Learning Models on FPGAs

A systematic review of FPGA-based DL deployment from a cross-layer perspective spanning model, compiler, architecture, runtime, and electronic design automation (EDA) is presented, highlighting that reliable cross-study comparison requires careful consideration of model configuration, precision, execution phase, batch...

Shuo Wang, Lei Chen, Chunsheng Tian et al. · 0 citations
Open access Aug 2026

A Compiler-Aware Framework for Partitioned Neural Network Inference on FPGA DPUs

This work analyzes the Vitis AI compiler and proposes an XIR-level splitting framework that generates independently compilable .xmodel fragments while preserving the context required for DPU mapping, and restores correct DPU mapping by addressing boundary-context loss and incomplete dependency collection.

Federico Buccellato, Luca Mannini, C. De Sio · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.