Event-Guided Online Video Super-Resolution
Abstract
Event-guided video super-resolution (VSR) leverages high-temporal-resolution event streams to address motion blur, rapid dynamics, and poor illumination that challenge frame-only VSR methods. However, most existing approaches emphasize reconstruction quality while overlooking real-time performance and computational efficiency, limiting their deployment in latency-sensitive scenarios. To overcome these issues, we present E2VSR, a lightweight and Efficient Event-guided VSR framework tailored for real-time applications. Operating under a causal setting with only current and past observations, E2VSR is designed for low-latency event-guided VSR. We propose an event-confidence adaptive propagation strategy comprising two key modules: the Event-induced Feature Modulation (EvFM) block for robust cross-modal event-frame integration, and the Event-Confidence Feature Fusion (EvCFF) block, which exploits events as motion cues for adaptive inter-frame aggregation. This design improves motion-aware temporal aggregation in challenging dynamic conditions, where event cues may provide complementary temporal information. Furthermore, an Implicit Event Reconstruction (IER) technique leverages event information during training to enrich feature representations without adding inference-time cost, enhancing spatial and temporal fidelity. Experimental results demonstrate that E2VSR achieves superior quantitative and qualitative performance while maintaining a low parameter count and computational cost.