Training-free online video anomaly detection (VAD) with frozen multimodal language models faces two coupled challenges: extracting reliable current-window semantics under causal and computational constraints, and maintaining temporal continuity without repeatedly transmitting high-dimensional history. Encoding history...
Ji Wang, Shuang-Qing Zhang, Guo-Sen Xie et al.· 0 citations
Weakly supervised video anomaly detection (WS-VAD) presents a significant challenge in security video surveillance, as it aims to accurately identify anomaly frames in untrimmed videos with only video-level supervision. Several recent studies exploit vision-language pre-training models, e.g., CLIP, to take advantage of...
Shuang-Qing Zhang, Wei Xu, Yu-Qi Fang et al.· IEEE Transactions on Informa...· 2 citations
Probe-VAD is proposed, an ordinal binary-probing framework that directly probes severity preferences from a frozen VLM, providing a simple interface for translating frozen VLM visual understanding into continuous, rank-sensitive anomaly scores without task-specific training or caption-based compression.
Jia-Wei Gu, Qi-Lin Zhao, Teng-Kuo Guo et al.· 1 citation
This work proposes a novel Text-Driven Video Anomaly Detection (TD-VAD) approach, which utilizes video-like text descriptions with temporal characteristics generated by LLM to train a VAD model, without any reliance on target-domain anomaly data.
Shuang-Qing Zhang, Lei-Lei Ma, Zhao Wang et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.