Multimodal deception detection via feature reconstruction and spatio-temporal consistency modeling
Abstract
Deception detection is an important task in security, forensic analysis, and human-computer interaction. Despite its significance, traditional unimodal approaches often suffer from limited representation capacity. To bridge this gap, this paper introduces a novel multimodal deception detection framework centered on feature reconstruction and temporal consistency modeling. Our method integrates textual, acoustic, and visual modalities by leveraging state-of-the-art pretrained backbones for robust feature extraction. To capture the long-range temporal dynamics inherent in deceptive behaviors, an LSTM-based encoder is incorporated. Furthermore, we design an AutoEncoder-based reconstruction mechanism to regularize the fusion process; by enforcing reconstruction constraints, the model is enabled to learn more compact, discriminative, and noise-resilient multimodal representations. Extensive experiments on the MDPE dataset demonstrate that our proposed framework outperforms baseline models across multiple metrics, including Accuracy, F1-score, and AUC. Ablation studies validate that the integration of feature reconstruction and temporal modeling provides a superior solution for identifying complex deceptive patterns.