Skip to content

Hardware-Aware Learned Representation Compression for Distributed In-Sensor Vision

Sep 2026 · 0 citations · 48 references
Computer Science

TL;DR

OASIS, a distributed in-sensor vision framework that uses a lightweight encoder to generate compact, task-relevant representations before off-chip transmission, is presented, demonstrating a practical hardware-algorithm co-design path for communication-efficient in-sensor vision.

Abstract

In-sensor computing reduces the cost of transmitting high-resolution image data by performing early-stage processing near the sensor. However, the logic chip integrated with a CMOS image sensor (CIS) is tightly constrained in compute and memory, limiting conventional deep neural network partitioning. We present OASIS, a distributed in-sensor vision framework that uses a lightweight encoder to generate compact, task-relevant representations before off-chip transmission. The encoder is trained end-to-end using task, entropy, and reconstruction objectives, while the decoder is used only during training. OASIS supports two complementary deployment paths. The first applies 4-bit quantization and Huffman coding while preserving the spatial structure required by classification and dense-prediction tasks. The second uses Sobol-based hyperdimensional computing (HDC) to transform the encoder latent into a fixed-dimensional binary hypervector for associative-memory classification. For the SwinViT-based VWW model, mapping a $3\times3\times8$ latent to a 64-dimensional hypervector provides an additional $1.77\times$ communication reduction with less than one percentage point of accuracy loss relative to the 128-dimensional configuration, yielding an overall $18{,}816\times$ reduction compared with raw 8-bit image transmission. We implement the digital near-sensor pipeline on an AMD Xilinx Zynq UltraScale+ FPGA and characterize it using direct board-level power measurements and Vivado post-implementation analysis, together with circuit-simulated CIS models and a 7-nm ASIC projection. Across visual wake-word classification, hand tracking, and eye tracking, OASIS reduces total system energy by approximately $2\times$-$4.5\times$ while maintaining competitive accuracy, demonstrating a practical hardware-algorithm co-design path for communication-efficient in-sensor vision.

View source

Similar papers

Conference Open access 2026

Key Algorithms of Convolutional Neural Networks and Hardware Implementation of Image Processing

Low-level hardware acceleration strategies to deconstruct the mapping from algorithm logic to silicon substrates are reviewed to provide strong guidelines for the hardware-software co-design of emerging ultra-low power edge Artificial Intelligence (AI) chips.

Linenxu Zhang · 0 citations
#artificial intelligence Preprint Sep 2026

Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer

This work builds a GPU library that emulates SC matrix multiplication at scale, exposes stream lengths as first-class kernel arguments, and evaluates SC end-to-end on image classification, object detection and instance segmentation, class-conditional image generation, and visual world-model planning.

Hao-Ran Jin, Kang-Qi Zhang, Ji-Rong Yang et al. · 0 citations
#machine learning Preprint Aug 2026

Distributed Semantic Segmentation With Improved Rate-Distortion Trade-Off

The effectiveness of the proposed source codecs are demonstrated by achieving state-of-the-art performance in distributed semantic segmentation at below 0.2 bits per pixel, measured using the mean intersection-over-union metric on ADE20K (Cityscapes).

Danish Nazir, Timo Bartels, Thorsten Bagdonat et al. · 0 citations
Preprint Sep 2026

VQ-LIC: Shared Vector-Quantized Learned Image Compression on a Resource-Constrained FPGA

Learned image compression (LIC) is hard to deploy on severely resource-constrained FPGAs, since how fast it actually runs depends not just on arithmetic count, but also on memory traffic, imbalance between different operations, and how the hardware batches its work. We present VQ-LIC, an asymmetric edge-cloud codec in...

Muhammad Fahd Ibrahim Bhatti, Abdullah Bin Faisal, Ahsan Usman et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.