Skip to content
Preprint

CascadeLUT: Information-Ordered Streaming Inference for Bandwidth-Constrained FPGAs

Aug 2026 · 0 citations · 21 references
Computer Science

TL;DR

This work presents CascadeLUT, an information-structured inference framework organized around bandwidth constraints, where instead of buffering the full input, features are partitioned into ordered subsets and predictions are progressively refined as subsets arrive, enabling deterministic streaming inference without runtime branching.

Abstract

Mapping neural networks to FPGAs enables low-latency, energy-efficient inference, particularly for lookup table (LUT)-based models that eliminate multipliers and map directly to reconfigurable fabric. While prior work achieves high compute efficiency, it typically assumes full-sample availability, causing pipeline stalls in bandwidth-limited streaming scenarios. Here, the bottleneck shifts from computation to data movement, as large input transfers limit throughput and energy efficiency. We present CascadeLUT, an information-structured inference framework organized around bandwidth constraints. Instead of buffering the full input, features are partitioned into ordered subsets and predictions are progressively refined as subsets arrive. The cascade statically controls which layers consume incoming features, enabling deterministic streaming inference without runtime branching. By co-designing feature scheduling with hardware dataflow, CascadeLUT reduces data movement while maintaining accuracy. Across datasets, it achieves 4.0 to 12.5 times lower latency, 3.0 to 5.0 times higher throughput and up to 13.8 times lower energy/sample than prior LUT baselines, using 1.2 to 4.4 times the LUTs of the smallest DWN baseline per task. We also demonstrate on-device input quantization integrated with LUT-based inference and present end-to-end FPGA results on real-world workloads, with 5 times reductions in quantization overhead.

View source

Similar papers

Open access Oct 2026

QUINN: Quantized Inference and Dataflow Characterization for Neural Networks on RISC-V Near-Memory Architectures

Machine learning inference is moving to the edge due to latency, energy, and privacy constraints. The growing complexity of edge workloads places increasing pressure on memory hierarchies, making data movement a major limitation for both performance and energy efficiency. Fixed-function accelerators inherit the memory...

Vincenzo Petrolo, Flavia Guella, Guido Masera et al. · 0 citations
Preprint Sep 2026

LLM Inference on IMC-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism

LLM inference has become an essential service, yet it imposes unprecedented demands on memory bandwidth, computational density, and communication efficiency. While IMC is a promising solution to the memory wall issue, the heterogeneous data dynamicity of LLM requires complementary resources to handle intermediate data...

Yimin Wang, Yue Jiet Chong, Xuan-Yao Fong · 0 citations
Aug 2026

Learning to schedule: interference-aware optimization for edge AI inference on shared GPUs

A time-varying integer program to minimize the long-term total cost of the edge AI inference system, including the inference latency, the inference error rate, the query-dispatching communication cost, and the energy consumption, subject to resource and workload constraints is proposed.

Ming-Tao Ji, Hehan Zhao, Lei Jiao et al. · 0 citations
Conference Aug 2026

Designing and Building an FPGA Accelerator That Uses Less Energy for DNN Inference

Deep Neural Networks (DNNs) are critical to modern AI applications, yet their deployment on standard CPUs and GPUs is constrained by high power consumption and computational latency, particularly in resource-constrained edge environments. To address these limitations, this paper presents the design and implementation o...

P. V. G. K. Rao, Dudekula Raziya · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.