This study introduces Vision Mamba (ViM), leveraging a Selective State Space Model to capture long-range temporal dependencies with linear computational complexity, validating the ViM as a highly efficient solution that balances computational feasibility with good performance in detecting complex criminal activities at the frame level.
Abstract
Video Anomaly Detection (VAD) faces persistent challenges, including annotated data availability, contextual dependency, and elevated false alarm rates. While Self-Supervised Learning (SSL) effectively mitigates label scarcity, the current Self-Supervised Multi-Task Learning (SSMTL) framework encounters two primary challenges. The first challenge pertains to the susceptibility of Conv3D-based encoders to overfitting. This study investigated several overfitting mitigation strategies integrated into the encoder architecture to address this issue. The second challenge concerns the quadratic computational costs inherent in Transformer-based encoders. As a solution, this study introduces Vision Mamba (ViM), leveraging a Selective State Space Model to capture long-range temporal dependencies with linear computational complexity. This efficiency is theoretically substantiated by an asymptotic time complexity analysis, demonstrating ViM’s superiority over Vision Transformers (ViT) regarding data depth dimensions and architectural depth layers. The comprehensive experiments yielded two key findings. First, the empirical results on UCSD Ped2 demonstrate that a 0.3 dropout rate provides superior stability for mitigating overfitting in Conv3D baselines compared to standard regularization. Second, evaluations on a theft-focused UCF-Crime subset confirm ViM as the most lightweight architecture, reducing Floating Point Operations (FLOPs) by a factor of 2-4 relative to alternatives. In terms of performance, ViM outperformed the Conv3D baseline (avg. +0.012) and rival VideoSwin (avg. -0.017), although it trails ViT in the AUC-ROC, Precision, and Recall metrics. Finally, this study validates the ViM as a highly efficient solution that balances computational feasibility with good performance in detecting complex criminal activities at the frame level.
This research introduces a novel memory-augmented CNN-ConvViT autoencoder framework for unsupervised video anomaly detection and introduces a Temporal-Aware Prototype Memory Module (TAPMM) that explicitly learns normal spatio-temporal behavior patterns.
Zero-shot industrial anomaly detection (ZIAD) aims to develop a unified model capable of directly identifying unseen anomaly categories in images without requiring reference samples. Recently, large-scale Vision-Language Models (VLMs) such as CLIP have shown great potential for solving this task. However, existing meth...
Tiyu Fang, Lin Zhang, Ran Song et al.· IEEE Transactions on Automat...· 0 citations
Weakly supervised video anomaly detection remains a challenging problem, primarily due to the scarcity of abnormal training samples and the lack of diverse feature representations, which hamper the learning of discriminative models. To address these issues, we introduce a novel weakly supervised cross-domain framework...
Mao-Wen Zhou, Erma Rahayu Mohd Faizal Abdullah, Aznul Qalid Md Sabri et al.· PLoS ONE· 0 citations
Deception detection is an important task in security, forensic analysis, and human-computer interaction. Despite its significance, traditional unimodal approaches often suffer from limited representation capacity. To bridge this gap, this paper introduces a novel multimodal deception detection framework centered on fea...
A meta-learning framework that combines Model-Agnostic Meta-Learning (MAML) with a dual-memory, transformer-based architecture and a dual memory backbone is proposed, providing a useful proof of concept where MAML has been shown to learn generalized anomaly and non-anomaly representations with a transformer based archi...
Shradha Mahadev Naik, Suja Palaniswamy, Nicola Conci· Journal of King Saud Univers...· 0 citations
The rapid growth of Generative Adversarial Networks (GANs) has made synthetic media incredibly realistic and led to the emergence of deepfakes that can harm cybersecurity, trust in the digital world, and information integrity. The current state of the art in deepfake detection faces several issues, such as limited gene...
Abdullah Alhejaili, Muhammad Binsawad· Electronics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.