UniMoRe: Unified Motion Representation via Gait IMU Pretraining
Abstract
Pretraining for inertial measurement unit (IMU) signals is gaining traction. Several strategies are devised based on complex architectures, multimodal fusion, or large multiactivity corpora. Yet, it remains unclear whether a simple, single-modality model pretrained on a narrowly focused motion domain can yield general-purpose motion representations. We introduce UniMoRe: unified motion representation via gait IMU pretraining, a transformer-based masked autoencoder (MAE) initialized from a pretrained masked EEG signal modeling autoencoder and trained exclusively on gait sequences from a single IMU sensor. By reconstructing masked inputs, UniMoRe learns rich temporal features that extend beyond gait. Despite its generic gait-centric pretraining, it generalizes effectively to diverse downstream tasks, including general action recognition, gait anomaly detection, and gait recognition, demonstrating that gait pretraining provides a transferable foundation for broader human activity understanding. Experiments show that UniMoRe is also resilient to perturbations—like masking, scaling, shifting, and rotation—highlighting its stable motion primitives learning rather than dataset-specific cues. Compared against a state-of-the-art IMU pretraining framework, UniMoRe achieves competitive or superior performance while using a simpler pretraining strategy and significantly narrower pretraining distribution. Our ablation studies establish the potential of masked reconstruction-based single-modality, single-sensor, and gait-specific IMU pretraining for unified and generalizable motion representations.