Skip to content
Preprint

MASQ: Mask-Aware Spatiotemporal Quantization for Unsupervised Skeleton Action Segmentation

Aug 2026 · 0 citations
Computer Science

TL;DR

A novel Mask-aware Action Spatiotemporal Quantization framework that decouples the conflicting tasks of spatial feature inference and temporal smoothing, and establishes a comprehensive and substantial leading advantage in the Mean over Frames accuracy.

Abstract

Unsupervised skeleton-based temporal action segmentation is a crucial task for understanding human behavior in long untrimmed sequences. Recent approaches often rely on discrete quantization to discover action boundaries from motion representations. However, when spatial masking is introduced for representation learning, it can introduce representation ambiguity, while discrete quantization further amplifies small fluctuations in the latent space. The interaction between these two factors often leads to unstable code switching and severe temporal jitter near action boundaries.To address these limitations, we propose a novel Mask-aware Action Spatiotemporal Quantization (MASQ) framework. Our framework decouples the conflicting tasks of spatial feature inference and temporal smoothing.In the spatial dimension, we introduce a Joint-Level Structured Dropout (JLSD) mechanism that masks the entire temporal trajectory of selected joints, to encourage the model to learn discriminative inter-joint coordination patterns. In the temporal dimension, we design a mask-aware velocity loss that enforces motion consistency only on visible joints, that prevents gradient conflicts caused by masked signals and stabilizing temporal predictions. Extensive experiments on three widely used skeleton datasets, including HuGaDB, LARa, and BABEL, demonstrate that the proposed MASQ framework significantly outperforms existing state-of-the-art unsupervised methods. In particular, our model establishes a comprehensive and substantial leading advantage in the Mean over Frames accuracy.

View source

Similar papers

Preprint Aug 2026

See the Change, Keep the Flow: Unsupervised Action Segmentation via Spectral-Temporal Representation Learning

Unsupervised action segmentation aims to discover latent action categories and their temporal organization without action annotations. Optimal transport-based methods provide structured frame-to-action assignments, however, their pseudo-label quality is fundamentally conditioned on the representation space used to cons...

Yun Li, Jun Xiao, Cong Zhang et al. · 0 citations
Preprint Sep 2026

MotionSpaceFlow: Representation-Aware Flow Matching in Direct Motion Space

Recent advances in diffusion and flow models have substantially improved text-driven human motion generation. Yet most methods generate in low-dimensional, temporally downsampled latent spaces learned primarily for reconstruction, a bottleneck that can limit generation quality and preclude direct manipulation of indivi...

Yu Qing, Kent Fujiwara · 0 citations
Preprint Aug 2026

Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining

This work introduces a weakly-supervised vision-language pretraining mechanism that transitions pooling kernels from pretraining, which aggregates skeleton features at the video level and aligns them with each video's known action text embeddings, to the inference phase that computes instance-level features without tra...

Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma et al. · 0 citations
Open access 2026

Cross-Modal Temporal Alignment for Action Grounding in Videos

A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.

Gengtian Shi, Chen-Hao Wu, Shao-Fei Wang et al. · 0 citations
Open access Sep 2026

Beyond Structural Symmetry: Elastic Spatio-Temporal Fluid Graph Convolutional Networks for Skeleton-Based Action Recognition

Skeleton-based human action recognition is a pivotal research area in computer vision. While conventional methods predominantly model human actions as rigid displacements strictly adhering to structural symmetry, the essence of authentic action is actually an elastic spatio-temporal fluid driven by highly asymmetric mo...

Xiang Yu, Cheng-Ming Xie, Ya-Qi Chen et al. · 0 citations
Open access Sep 2026

DiMoP: Diffusion-Driven Motion Representation Learning With Frame-Level Pseudo-Classification for Skeleton-Based Action Recognition

Robust skeleton-based action recognition requires representations that capture a wide spectrum of motions, from subtle to moderate and strong ones. Existing methods often focus on strong motions. This paper introduces DiMoP, a masking- and diffusion-driven motion representation learning method with frame-level pseudo-c...

Shanaka Ramesh Gunasekara, Wan-Qing Li, Nikalal Kaldera et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.