Skip to content

Spike2Rep: Learning Compact Spatiotemporal Representations for Recognition With High-Speed Spike-Based Image Sensor

Aug 2026 · IEEE Sensors Journal · Vol 26, pp. 25171-25180 · 0 citations · 43 references

Abstract

High-speed spike-based image sensors have shown great potential for high-speed object recognition due to their ultrahigh temporal resolution and static scene retention. To fully exploit this hardware capability, converting spike streams into structured representations suitable for deep neural networks remains a key challenge. Notably, object categories are temporally invariant, while spike streams exhibit rich temporal dynamics, suggesting exploitable spatiotemporal redundancy. In this work, Spike2Rep is proposed to exploit the intrinsic spatiotemporal redundancy in spike streams for a compact representation. In particular, the temporal dynamics are treated as motion cues to reveal spatial importance, and temporal redundancy is used for adaptive temporal aggregation. To implement this, the discrete wavelet transform (DWT) is applied to obtain a compact time–frequency representation from spike streams. Subsequently, the temporal difference is employed to capture motion cues and generate spatial importance maps, which are added to enhance features. Furthermore, temporal and channel weights are predicted via a linear layer, which is used for adaptive modulation. Finally, the temporal averaging is adopted to produce a compact spatiotemporal representation. On the S-CALTECH dataset, the proposed method achieves up to 81.27% top-1 accuracy using 2.2 ms of spike streams and 87.73% with 6.4 ms, surpassing the prior state-of-the-art (SOTA) methods on this dataset.

View source