Mamba-Guided Diffusion for Prediction-Based Industrial Video Anomaly Detection
Abstract
Industrial video anomaly detection must handle more than one failure mode: motion, position, rhythm, and object state may all drift We present MG-DDAD, a prediction-based normal-future anomaly detector built on a re-implementation of the DiffiiMa predictor. Given 16 frames of RGB history, it predicts a gap-4 normal-future frame at 128 × 128. Mamba supplies long-range temporal context; a DiT / diffusion residual path refines spatial detail. The contribution is ValNorm (validation-normalized multi-branch scoring). It turns prediction behavior into frame-level evidence by separating expected-versus-observed motion discrepancy from local prediction inconsistency, and keeps appearance and object-state cues as diagnostic branches. This keeps model-based evidence apart from no-model motion priors. We evaluate on IPAD, with 12 synthetic and 4 real industrial scenes, using per-scene training and frame-level AUC. Under the same prediction-based protocol, overall AUC moves from 69.4% for a model-free baseline to 72.1% with the Mamba predictor and 75.4% with full MG-DDAD. The ordering holds across scenes, with full-model means of 77.1% on synthetic and 70.2% on real scenes. Because this protocol differs from the standard IPAD benchmark, these results are for controlled ablation and reference, not direct external comparison.