Reconstructive Visual Tuning for Weakly Supervised Video Anomaly Detection
Abstract
Weakly supervised video anomaly detection (WS-VAD) presents a significant challenge in security video surveillance, as it aims to accurately identify anomaly frames in untrimmed videos with only video-level supervision. Several recent studies exploit vision-language pre-training models, e.g., CLIP, to take advantage of cross-modal relationships by adapting the pre-trained vision-language associations to the WS-VAD task. However, the supervision in these methods is derived solely from the coarse video-level labels. Consequently, they lack the precise supervision to understand video details comprehensively and exhibit systematic visual shortcomings, such as failure to recognize fine-grained anomaly patterns that require rich and detailed information. To address this problem, we propose a novel Reconstructive Visual Tuning (RVT) approach for the WS-VAD task. Instead of the naive exploitation of coarse video-level supervision, we enhance the WS-VAD model by reconstructing input video features to regularize the visual representations. By doing so, it capitalizes on the inherent detail and richness within the input features themselves, which are typically overlooked in previous WS-VAD approaches. Specifically, RVT employs a reconstruction objective that operates on the input video features, thereby circumventing the spatial redundancy inherent in raw RGB value regression. Furthermore, to better capture both detailed local and global temporal patterns, we design a global-local collaborative temporal modeling module (GLCT) to capture temporal dependencies from global and local perspectives. Extensive experiments demonstrate that the proposed RVT model delivers consistent and significant performance improvements across various VAD datasets. In comparison with existing state-of-the-art methods, RVT achieves competitive performance, underscoring the effectiveness of the proposed approach.