Skip to content
Open access

Meta-learning guided weakly supervised video anomaly detection with dual memory and temporal attention

Aug 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 42 references

TL;DR

A meta-learning framework that combines Model-Agnostic Meta-Learning (MAML) with a dual-memory, transformer-based architecture and a dual memory backbone is proposed, providing a useful proof of concept where MAML has been shown to learn generalized anomaly and non-anomaly representations with a transformer based architecture and a dual memory backbone.

Abstract

Weakly supervised video anomaly detection (WSAD) aims to localise anomalous events in untrimmed videos using only video-level labels. Existing multiple instance learning (MIL) methods often suffer from poor generalisation to unseen anomaly types, unstable temporal attention, and limited adaptability when only a few labelled examples are available. To address these challenges, we propose a meta-learning framework that combines Model-Agnostic Meta-Learning (MAML) with a dual-memory, transformer-based architecture. The model incorporates a dual-branch temporal attention module that captures both long-range semantic dependencies and local temporal proximity, separate memory banks for normal and abnormal prototypes with gated inhibition, metric-learning constraints, and variational latent regularisation. MAML explicitly trains the model for rapid adaptation across heterogeneous anomaly distributions, forcing it to acquire task-invariant representations rather than memorising static training statistics. Extensive experiments on two standard benchmarks yield competitive frame-level AUC of 93.60% on XD-Violence and 86.10% on UCF-Crime. One of our main contributions is the demonstration of very good metrics for zero and few-shot cross dataset transfer experiments, using only a handful of weakly labelled videos. We thus provide a useful proof of concept where MAML has been shown to learn generalized anomaly and non-anomaly representations with a transformer based architecture and a dual memory backbone. A t-SNE analysis of the memory prototypes confirms that MAML produces well-separated normal and abnormal clusters, while without meta-learning the memory banks collapse into entangled representations. The model is also shown to be computationally efficient, confirming its practical value for real-world surveillance deployment.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection

An adaptive temporal modeling framework for WSVAD that explicitly accounts for variations in video dynamics across multiple temporal granularities is proposed and an adaptive similarity-based fusion strategy that dynamically integrates anomaly scores into video-level predictions is proposed, replacing fixed top-k aggre...

Chang-Yi Li, Yu Xiao · 0 citations
Preprint Aug 2026

MuST-VAD: Mutual Structured Learning for Video Anomaly Detection

In this paper, we propose MuST-VAD, a mutual structured learning framework for weakly supervised video anomaly detection (VAD) in which an anomaly detector and a large vision-language model (LVLM) exchange their acquired knowledge. Detectors in weakly supervised VAD learn anomaly scores from features extracted by a fix...

Satoshi Hashimoto, Hitoshi Nishimura, Mori Kurokawa · 0 citations
#artificial intelligence Preprint Oct 2026

Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but of...

Mohd Ubaid Wani, Sara Atito, Josef Kittler et al. · 0 citations
Aug 2026

TASG-VAD: Weakly Supervised Video Anomaly Detection via Temporal Variation Attention and Adaptive Saliency Guidance

Weakly supervised video anomaly detection (WSVAD) is important in intelligent surveillance. Existing methods often overemphasize salient abnormal segments, overlook subtle clues, and model temporal dependencies ineffectively. To address these issues, we propose TASG-VAD, an efficient anomaly detection framework. The pr...

Lihu Pan, Mingkai Hu, Lin-Liang Zhang et al. · 0 citations
2026

Sparsity-Controllable Normality Learning With Vision–Language Models for Scenario-Related Video Anomaly Detection

Video anomaly detection (VAD) is critical for automation systems and security surveillance. Recently, multimodal vision–language models (MLLMs) have attracted increasing attention due to their rich pre-trained knowledge and strong explainability. However, existing MLLM-based approaches struggle to adapt to real-world s...

Jiangyun Chen, Yuanjie Dang, Peng Chen et al. · 0 citations
2026

Reconstructive Visual Tuning for Weakly Supervised Video Anomaly Detection

Weakly supervised video anomaly detection (WS-VAD) presents a significant challenge in security video surveillance, as it aims to accurately identify anomaly frames in untrimmed videos with only video-level supervision. Several recent studies exploit vision-language pre-training models, e.g., CLIP, to take advantage of...

Shuang-Qing Zhang, Wei Xu, Yu-Qi Fang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.