Recognition and Modeling of User Interaction Behaviors in Virtual Reality Based on 3D Convolutional Neural Networks
Abstract
Accurate recognition of continuous user interactions in virtual reality (VR) is essential for intelligent perception systems and provides methodological support for multimodal information processing in electromagnetic sensing and nextgeneration human–machine interaction. However, fragmented action sequences, weak long-range dependencies, and semantic discontinuities remain major challenges in complex immersive environments. This study proposes a crosssegment spatiotemporal attention and behavior pattern modeling framework based on a three-dimensional convolutional neural network (3D CNN). The framework integrates multi-layer 3D convolution with residual learning for robust feature extraction, employs cross-segment spatiotemporal attention and Bidirectional Long Short-Term Memory (BiLSTM) networks to capture long-range temporal dependencies, and introduces graph neural networks with boundary smoothing and graph consistency constraints to preserve behavioral structure and action continuity. Experimental results demonstrate an average recognition accuracy of 87.2% for complex actions, a behavior pattern consistency of 95.2%, and real-time inference performance of 39.2 FPS while maintaining superior robustness under occlusion and motion blur. Beyond VR interaction analysis, the proposed hierarchical spatiotemporal representation and structured modeling strategy provides useful insights for adaptive signal interpretation, multimodal perception, and intelligent decisionmaking in electromagnetic sensing and wireless interactive systems.