Skip to content
Open access

Vision Mamba With Joint Spatiotemporal Features for Efficient Video Representation Learning in Self-Supervised Scheme

2026 · IEEE Access · Vol 14, pp. 132266-132281 · 0 citations · 58 references

TL;DR

This study introduces Vision Mamba (ViM), leveraging a Selective State Space Model to capture long-range temporal dependencies with linear computational complexity, validating the ViM as a highly efficient solution that balances computational feasibility with good performance in detecting complex criminal activities at the frame level.

Abstract

Video Anomaly Detection (VAD) faces persistent challenges, including annotated data availability, contextual dependency, and elevated false alarm rates. While Self-Supervised Learning (SSL) effectively mitigates label scarcity, the current Self-Supervised Multi-Task Learning (SSMTL) framework encounters two primary challenges. The first challenge pertains to the susceptibility of Conv3D-based encoders to overfitting. This study investigated several overfitting mitigation strategies integrated into the encoder architecture to address this issue. The second challenge concerns the quadratic computational costs inherent in Transformer-based encoders. As a solution, this study introduces Vision Mamba (ViM), leveraging a Selective State Space Model to capture long-range temporal dependencies with linear computational complexity. This efficiency is theoretically substantiated by an asymptotic time complexity analysis, demonstrating ViM’s superiority over Vision Transformers (ViT) regarding data depth dimensions and architectural depth layers. The comprehensive experiments yielded two key findings. First, the empirical results on UCSD Ped2 demonstrate that a 0.3 dropout rate provides superior stability for mitigating overfitting in Conv3D baselines compared to standard regularization. Second, evaluations on a theft-focused UCF-Crime subset confirm ViM as the most lightweight architecture, reducing Floating Point Operations (FLOPs) by a factor of 2-4 relative to alternatives. In terms of performance, ViM outperformed the Conv3D baseline (avg. +0.012) and rival VideoSwin (avg. -0.017), although it trails ViT in the AUC-ROC, Precision, and Recall metrics. Finally, this study validates the ViM as a highly efficient solution that balances computational feasibility with good performance in detecting complex criminal activities at the frame level.

Read PDF

Similar papers

Sep 2026

Spatio-temporal video anomaly detection via CNN-ViT autoencoder and farneback optical flow

This research introduces a novel memory-augmented CNN-ConvViT autoencoder framework for unsupervised video anomaly detection and introduces a Temporal-Aware Prototype Memory Module (TAPMM) that explicitly learns normal spatio-temporal behavior patterns.

Vandana Pathak, Manoj Diwakar, Neeraj Kumar Pandey et al. · 0 citations
2026

Cross-Modal Guidance Learning for Zero-Shot Industrial Anomaly Detection

Zero-shot industrial anomaly detection (ZIAD) aims to develop a unified model capable of directly identifying unseen anomaly categories in images without requiring reference samples. Recently, large-scale Vision-Language Models (VLMs) such as CLIP have shown great potential for solving this task. However, existing meth...

Tiyu Fang, Lin Zhang, Ran Song et al. · 0 citations
Open access Sep 2026

Learn the interactions: Weakly supervised video anomaly detection with human-object interactions

Weakly supervised video anomaly detection remains a challenging problem, primarily due to the scarcity of abnormal training samples and the lack of diverse feature representations, which hamper the learning of discriminative models. To address these issues, we introduce a novel weakly supervised cross-domain framework...

Mao-Wen Zhou, Erma Rahayu Mohd Faizal Abdullah, Aznul Qalid Md Sabri et al. · 0 citations
Conference Sep 2026

Multimodal deception detection via feature reconstruction and spatio-temporal consistency modeling

Deception detection is an important task in security, forensic analysis, and human-computer interaction. Despite its significance, traditional unimodal approaches often suffer from limited representation capacity. To bridge this gap, this paper introduces a novel multimodal deception detection framework centered on fea...

Yao-Dong Zhou, Zong-Yu Zhang, Chun-Hua Ren · 0 citations
Open access Aug 2026

Meta-learning guided weakly supervised video anomaly detection with dual memory and temporal attention

A meta-learning framework that combines Model-Agnostic Meta-Learning (MAML) with a dual-memory, transformer-based architecture and a dual memory backbone is proposed, providing a useful proof of concept where MAML has been shown to learn generalized anomaly and non-anomaly representations with a transformer based archi...

Shradha Mahadev Naik, Suja Palaniswamy, Nicola Conci · 0 citations
Open access Sep 2026

Explainable Spatio-Temporal Attention Framework for GAN-Based Deepfake Detection

The rapid growth of Generative Adversarial Networks (GANs) has made synthetic media incredibly realistic and led to the emergence of deepfakes that can harm cybersecurity, trust in the digital world, and information integrity. The current state of the art in deepfake detection faces several issues, such as limited gene...

Abdullah Alhejaili, Muhammad Binsawad · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.