Skip to content
#edge computing Preprint

LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices

Aug 2026 · 0 citations · 76 references
Computer Science Engineering

TL;DR

A compact and configurable event-driven autoencoder that efficiently compresses neuromorphic data while preserving essential spatiotemporal structure for downstream inference and demonstrates the potential of compact event-driven models to advance environmentally conscious, low-power AI systems for high-speed perception in autonomous, mobile, and embedded computing environments.

Abstract

Event-based vision has emerged as a promising paradigm for energy-aware artificial intelligence (AI), offering sparse, low-latency visual signals that reduce redundant data processing and support sustainable edge computing. However, the asynchronous and noise-prone nature of event streams creates challenges for conventional deep learning models, which are often too computationally intensive for low-power embedded platforms. This work presents a compact and configurable event-driven autoencoder that efficiently compresses neuromorphic data while preserving essential spatiotemporal structure for downstream inference. The architecture integrates lightweight convolutional encoding with robust performance under adaptive event thresholding and a minimal classifier head, enabling substantial reductions in computational cost without degrading recognition fidelity. Extensive evaluations on the Smart Event Face Dataset (SEFD) and Event-Based Crossing Dataset (EBCD) show that the proposed framework achieves competitive or superior accuracy compared to YOLOv9 while requiring up to 35.6$\times$ fewer parameters. To assess real-world sustainability, the model is deployed on resource-constrained hardware: a Raspberry Pi 4B and a NVIDIA Jetson Nano. On NVIDIA Jetson Nano, it delivers real-time throughput of 44.8 FPS. On a Raspberry Pi 4B CPU, the 50\% autoencoder classifier consumes 16.19 J for the evaluated inference workload, corresponding to approximately 726.3$\times$ lower energy consumption than YOLOv9 under the same evaluation protocol. These results demonstrate the potential of compact event-driven models to advance environmentally conscious, low-power AI systems for high-speed perception in autonomous, mobile, and embedded computing environments.

View source

Similar papers

Conference Open access Jul 2026

Energy-Efficient Context-Aware Multimodal AI Inference at the Edge

Deploying multimodal artificial intelligence models on resource-constrained edge devices faces inherent bottlenecks in energy consumption and computational latency, as conventional full-modality inference pipelines keep all perception encoders active regardless of task context and environmental conditions, causing substantial power waste and degraded real-time performance. To address this challenge, this work presents an energy-efficient context-aware multimodal edge inference framework featuring a lightweight modality activation sparsity evaluation unit, dynamic computation path scheduling, and a cross-modal speculative skipping mechanism. The framework quantifies the information value of each input modality in real time according to scene context, task complexity, and device power status, and adaptively activates or deactivates corresponding visual, audio, and sensor encoders, while tuning model quantization precision and operator fusion strategies to align with runtime resource budgets. Validated on NVIDIA Jetson Nano and Raspberry Pi 5 edge platforms across VQAv2, MMBench, and multimodal perception benchmarks, the proposed framework delivers a 42.3% reduction in end-to-end energy consumption and a 30%-65% decrease in inference latency against static full-modality baselines, alongside $\mathbf{1. 5} \times$ to $\mathbf{2. 3} \times$ higher throughput with an accuracy loss no more than 1.2%. The runtime context scheduling module introduces less than 9 ms of latency overhead and only 0.32 W of additional power draw, with per-inference energy as low as 0.6 J; for battery-powered mobile edge devices, the framework extends continuous operating duration by over 72% under typical perception workloads. These findings confirm that context-aware adaptive scheduling can dramatically boost the energy efficiency of edge multimodal inference without sacrificing task performance, offering a viable deployment solution for low-power Internet of Things, intelligent surveillance, and human-robot interaction scenarios.

Xiaotian Fang, Ya-Hui Yang, Shuyuan Wang · 0 citations
Aug 2026

DSF-Net: Dual-strategy fusion for efficient audio-visual sound event localization and detection.

Audio-visual sound event localization and detection (AVSELD) seeks to identify and locate sound-emitting objects by leveraging both audio and visual data. Current methods primarily rely on convolutional neural networks (CNNs), whose constrained receptive fields limit their ability to capture broader contextual information. Although Transformer-based architectures exhibit considerable proficiency in capturing global contextual information, their efficacy is impeded by the quadratic computational complexity associated with processing long-range dependencies. This poses a significant bottleneck, particularly in scenarios involving longer sequence lengths. To overcome this limitation, we propose DSF-Net, a novel neural network that introduces a dual-strategy fusion approach for the AVSELD task. Built upon an efficient state-space model backbone to ensure linear complexity, DSF-Net is designed for robust and computationally efficient multi-modal comprehension. The proposed dual strategies consist of: (1) an Adaptive Frequency Fusion module that aligns and integrates features in the frequency domain, and (2) an Audio-aware Aggregation module that performs advanced feature integration while considering the consistency between modalities. These strategies are embedded within a progressive fusion framework to enhance overall feature learning. Extensive experiments on the STARSS2023 dataset validate our dual-strategy approach, demonstrating that DSF-Net achieves state-of-the-art performance and outperforms existing methods. The source codes are publicly available at https://github.com/Devin-Pi/avseld-mamba.

Rendong Pi, Yingchao Zhang, Wei Rao et al. · 0 citations
Preprint Aug 2026

LITEWAY: LIghtweight HAR via Temporal Efficient highWAY

Wearable human activity recognition (HAR) remains challenging due to the computational and energy constraints of deep learning models on resource-limited devices. Existing lightweight approaches often rely on recurrent architectures (e.g., GRU and LSTM), limiting parallelism and increasing inference latency. We propose LITEWAY, a modality-agnostic, fully convolutional framework for multichannel sensor time series that replaces recurrent temporal modeling with structured convolutional decomposition. LITEWAY combines lightweight convolutional blocks, strided temporal processing, and convolution-attention pooling to efficiently capture temporal dependencies while reducing computational complexity. We evaluate LITEWAY on 16 HAR datasets against TinyHAR, TinierHAR, and MLP-HAR. LITEWAY achieves competitive macro F1 while reducing model size by 4.06x-9.52x (Light) and 3.87x-9.07x (Full) compared with TinyHAR and TinierHAR. Deployment experiments further show energy reductions of 2.29x-3.14x (Light) and 1.46x-2.01x (Full) compared with TinierHAR and MLP-HAR, highlighting efficient fully convolutional temporal modeling for wearable HAR. The source code is publicly available at https://github.com/dominique-nshimyimana/liteway.

Dominique Nshimyimana, V. F. Rey, Mengxi Liu et al. · 0 citations
Open access 2026

Lightweight Neuromorphic Perception: Porting Low-Latency Privacy-Responsive Human Motion Analysis to Constrained Edge Architectures

Neuromorphic vision sensors offer significant advantages for real-time embedded perception due to their microsecond latency, high temporal resolution, wide dynamic range, and low power consumption. However, deploying multi-stage event perception pipelines onto edge hardware remains fundamentally constrained by memory, throughput, and execution bottlenecks on resource-limited silicon. In this work, we introduce an end-to-end, real-time perception pipeline deployed on a Raspberry Pi 5 platform, establishing a hardware-aware reference baseline for future Neural Processing Unit (NPU) and neuromorphic architectures. Our system integrates a custom attention-enhanced YOLOv8-nano model for joint person and face detection, a parameter-free ByteTrack multi-object tracking framework, and a lightweight kinematic feature extraction engine for behavioral triage and velocity-based activity classification. To evaluate localized spatial trade-offs—particularly critical for in-cabin automotive application domains such as Driver and Occupant Monitoring Systems (DMS/OMS) we systematically quantify the sensor bias configurations and optical geometries (76° vs. 104° Fields of View) across detection accuracy, temporal consistency, face region extraction, and kinematic activity mapping. The detection backbones are fine-tuned on a custom indoor event dataset and exported using static ONNX execution graphs optimized for edge deployment. Operating strictly on sparse spatiotemporal event representations without intermediate RGB frame reconstruction, our pipeline maintains robust, privacy-responsive performance under extreme lighting and motion dynamics, sustaining 19–24 FPS end-to-end throughput. By providing complete quantitative and qualitative benchmarks, this framework establishes a reproducible baseline for low-latency edge-AI deployment in automotive cabin monitoring, autonomous robotics, and assistive healthcare. The project details, fine-tuned models, empirical results, and curated event datasets are publicly available at http://mali-farooq.github.io/NeuroVision

Muhammad Ali Farooq, G. Costache, Peter Corcoran · 0 citations
2026

ELKFANet: Efficient Large Kernel Network Enabled Frame Acquisition for Low SNR Scenarios

Frame acquisition is crucial to the entire receiving process. However, both traditional methods and existing deep learning approaches encounter performance challenges in low signal-to-noise ratio (SNR) scenarios. To address this issue, this letter proposes a novel deep learning-enabled frame acquisition scheme. We first introduce a data augmentation strategy that incorporates random time offsets and additive noise, specifically designed to mimic non-ideal reception conditions, such as capturing an incomplete preamble. Subsequently, we propose a frame acquisition neural network based on an efficient large kernel module, namely ELKFANet. This module decouples a standard large convolutional kernel into a sequence of horizontal, vertical, and pointwise convolutions, thereby capturing periodic and structural signal features through a global receptive field while maintaining computational efficiency. Experimental results show that under low SNR conditions ranging from −20 dB to 0 dB, ELKFANet achieves superior overall detection performance compared with existing methods across different channel models. Furthermore, the proposed module strikes an effective balance between computational complexity and detection performance, and exhibits strong scalability across different bandwidth scenarios.

Ming Zeng, Kui Xu, Xingyu Zhou et al. · 0 citations

Related blog posts

Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.