Skip to content
Open access

Deployment and Optimization of YOLOv8n Edge Object Detection on RISC-V K230

2026 · IEEE Access · Vol 14, pp. 138588-138602 · 0 citations · 29 references

Abstract

To address the challenges of deploying deep neural networks (DNNs) on heterogeneous edge platforms for industrial defect detection, this study optimizes the YOLOv8n model for the RISC-V K230 platform. Microbenchmarking and pipeline analysis reveal three primary bottlenecks impeding lightweight deployment: memory bandwidth limitations, computational granularity mismatch, and pipeline stalls. To address these bottlenecks, we propose HSA-YOLO, a hardware-spectrum-aware hierarchical optimization network. The model features a three-tier optimization architecture. At the shallow layer, we developed a hardware-decoupled KPU adaptation unit (HDKU) by aligning operator granularity with hardware parallelism, employing operator fusion, and reusing memory access. This approach reduces computational fragmentation and enhances feature expression for small targets, enhancing KPU utilization without compromising detection accuracy. At the middle layer, we introduced a differential precision scheduling (DPS) strategy, which allocates data bit widths based on path semantics, maximizing RVV throughput while preserving accuracy, thereby optimizing the accuracy-throughput trade-off under RVV bandwidth constraints. At the deep layer, we implemented a static cache residency module (SCRB), confining intermediate feature activations within the L2 cache and employing channel compression to reduce off-chip memory access and alleviate memory bandwidth pressure. Experimental results demonstrate that, compared to the baseline YOLOv8n, the optimized M3 model achieves an mAP@0.5 of 96.62% with an end-to-end latency of 33.59 ms. The complete HSA-YOLO model trades a small amount of accuracy and latency for an 11.3% reduction in parameters and improved memory access efficiency, ultimately achieving an mAP@0.5 of 95.50%, reducing end-to-end inference latency by 71.4%, and increasing throughput by $3.5\times $ , meeting the requirements of edge industrial detection scenarios with different constraints.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.