Aug 2026· Journal of Supercomputing· Vol 82· 1 citation· 52 references
TL;DR
A hardware algorithm co-optimization framework for high-throughput deployment of deep object detection models on reconfigurable architectures, demonstrating the effectiveness of algorithm architecture co-design in overcoming memory and parallel efficiency bottlenecks and providing a practical solution for real-time high-performance deployment of deep learning models.
A systematic review of FPGA-based DL deployment from a cross-layer perspective spanning model, compiler, architecture, runtime, and electronic design automation (EDA) is presented, highlighting that reliable cross-study comparison requires careful consideration of model configuration, precision, execution phase, batch...
Shuo Wang, Lei Chen, Chunsheng Tian et al.· Electronics· 0 citations
To address the challenges of deploying deep neural networks (DNNs) on heterogeneous edge platforms for industrial defect detection, this study optimizes the YOLOv8n model for the RISC-V K230 platform. Microbenchmarking and pipeline analysis reveal three primary bottlenecks impeding lightweight deployment: memory bandwi...
Shu Lan, Qiao-Yu Xu, Yue Sun· IEEE Access· 0 citations
Low-level hardware acceleration strategies to deconstruct the mapping from algorithm logic to silicon substrates are reviewed to provide strong guidelines for the hardware-software co-design of emerging ultra-low power edge Artificial Intelligence (AI) chips.
Linenxu Zhang· MATEC Web of Conferences· 0 citations
Traffic object detection for deployment is constrained by accuracy, latency, model size, and computational cost. Under a unified dataset, deployment pipeline, and single-GPU hardware platform (RTX 4090), this paper evaluates the combined effects of structured pruning and low-precision inference. Using YOLOv5s as the pr...
Hao-Jing Wang· ITM Web of Conferences· 0 citations