XRPoseSync: A Synchronized Multimodal Dataset for EdgeXR Pose Prediction
Abstract
Existing VR datasets lack synchronized multimodal pairing of headset tracking and egocentric visual context, making it difficult to systematically study when and how visual cues improve head-motion prediction. While prior prediction work suggests visual information can help, these claims often rest on isolated models rather than shared benchmarks with aligned pose, video, and external ground truth. We present XRPoseSync, a synchronized multimodal dataset with 65 dual-stream pose sessions, 57 of which also include first-person ego-view video captured from the SteamVR VR View mirror at the headset’s display refresh rate (effective ≈ 60 FPS), together with OpenXR device tracking (200 Hz), Qualisys motion-capture reference (500 Hz), segment metadata, and synchronization diagnostics for controlled multimodal analysis. The dataset is designed to expose realistic interactive VR motion under a common timeline rather than to privilege a single predictor or modality. We include a lightweight pose-history versus pose-plus-video benchmark based on dense optical-flow descriptors. The results show that visual cues do not uniformly help prediction: in this baseline, video gains are horizon- and entropy-dependent, with the most notable translation improvement of 0.48 cm at the 600 ms horizon, while gains are negligible at short horizons and at 800–1000 ms. These findings should be interpreted as evidence of research questions enabled by XRPoseSync, not as a final statement on the best visual representation for XR pose prediction. XRPoseSync provides synchronized pose-video pairs, external reference streams, alignment metadata, entropy-based characterization, and user-disjoint splits for reproducible multimodal XR analysis.