Skip to content

A High-Performance Neural Rendering Accelerator With Dual-Lane Micro-MLPs and Hierarchical Latency-Hiding Scheduling

Sep 2026 · IEEE Transactions on Circuits and Systems Part 1: Regular Papers · Vol 73, pp. 6129-6140 · 0 citations · 34 references

Abstract

While Neural Radiance Fields (NeRF) have transformed 3D vision, their prohibitive computational and memory demands restrict real-time deployment on power-constrained edge devices. To bridge this gap, we propose a novel neural rendering accelerator that orchestrates three architectural innovations to maximize throughput and energy efficiency. First, to resolve computation bottlenecks, we introduce a Dual-Lane Backend Mechanism utilizing Fused Micro-Multi-Layer Perceptrons (Micro-MLPs), boosting sample processing throughput by $1.98\times $ . To minimize memory traffic through data locality, we implement Multi-Resolution Spatial Partitioning Strategy with Adaptive Ray Clustering, effectively exploiting scene sparsity, and improving the cache hit rate up to 95%. Finally, to hide memory latency, a Hierarchical Scheduling Framework with Proactive Prefetching is employed, reducing execution stalls by 77%. Experimental results on a field-programmable gate array (FPGA) prototype demonstrate 94.7 frames per second (FPS) at $800\times 800$ resolution within a 6.4 W power consumption, achieving a 42.2% frame rate improvement over state-of-the-art FPGA implementations. Furthermore, application-specific integrated circuit (ASIC) synthesis results based on 28 nm CMOS technology projects a peak throughput of 440 FPS with a power consumption of 268 mW, translating to a highly competitive energy efficiency of 0.76 nJ/pixel. By delivering superior energy efficiency while maintaining high visual fidelity with a Peak Signal-to-Noise Ratio (PSNR) > 30 dB, this work paves the way for ubiquitous, photo-realistic rendering on the edge.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.