Skip to content
Preprint

Motion-Saliency Complementary Masked Modeling for Point Cloud Video Understanding

Aug 2026 · 0 citations · 43 references
Computer Science

TL;DR

MoSaiC, a novel Motion-Saliency Complementary masked modeling framework for self-supervised point cloud video representation learning, couples three components: Curriculum Motion-Saliency Masking (CMSM), which guides the masking process toward motion-salient tokens under a curriculum schedule; Normal-Flow Motion (NFM) modeling, which supervises the local rigid rotation of each token in the Lie algebra so(3) as an explicit geometric motion target.

Abstract

Point cloud video representation learning is crucial for 3D dynamic scene understanding. In this paper, we propose MoSaiC, a novel Motion-Saliency Complementary masked modeling framework for self-supervised point cloud video representation learning. MoSaiC couples three components: Curriculum Motion-Saliency Masking (CMSM), which guides the masking process toward motion-salient tokens under a curriculum schedule; Normal-Flow Motion (NFM) modeling, which supervises the local rigid rotation of each token in the Lie algebra so(3) as an explicit geometric motion target; and Cross-view Token Consistency Prediction (CTCP), which enforces consistency between two complementary masked views at the token level. Together, these components allow MoSaiC to effectively capture both appearance and motion dynamics. Extensive experiments on multiple downstream tasks, including action recognition, temporal action segmentation, and point-level semantic segmentation, demonstrate the effectiveness of our approach.

View source

Similar papers

Preprint Aug 2026

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation

This work introduces CamChoreo, a benchmark of 4,229 real single-shot clips with expert-annotated temporal segments, and proposes CamDistill, which distills the same geometric knowledge into lightweight camera tokens during training and removes the 3D model at inference.

Da-Zhao Du, Shiyan Du, Jian Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SWT: Self-Supervised Video Object Segmentation via Sliding, Wavelet and Transportation

Video Object Segmentation (VOS) aims to accurately segment target objects from consecutive video frames and track the changes of the objects in each frame of the video. Conventional VOS methods typically demand substantial quantities of pixel-level labeled video sequences for fully supervised learning, which limits the...

Zheng Zhu, Jia-Qing Fan, Han-Wen Qian et al. · 0 citations
Preprint Sep 2026

VideoTok4D: A 4D-Aware Video Tokenizer for Compact World Representation

Video tokenizers have emerged as a cornerstone of modern video modeling, underpinning progress in compression, reconstruction and generation by mapping high-dimensional visual signals into compact latent spaces. However, despite this progress, current tokenization paradigms largely remain within the 2D visual domain, t...

Xin-Yi Chen, Han-Xin Zhu, Xi-Jun Wang et al. · 0 citations
Open access Sep 2026

Vision–language guided semantic-geometric transformer for memory-efficient 3D scene understanding

Recent 3D Transformers have become a dominant framework for point-cloud segmentation by modeling spatial context in sparse 3D scenes. However, geometry and color alone provide limited high-level semantic cues, especially for cluttered boundary regions, visually similar objects, and long-tail categories. To address th...

Li-Cheng Liu, Yu Li, Fu-Yong Liu · 0 citations
Preprint Aug 2026

Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos

Self-supervised representation learning for 4D point cloud videos is challenging because annotations are costly and reconstruction-based pretraining can overemphasize low-level geometric details. We propose a JEPA-style framework that learns from unlabeled spatiotemporal point clouds through latent point-tube predictio...

Jheng-Ling Lee, Shangsheng Chen · 0 citations
Preprint Aug 2026

Mask What Matters: Saliency-Guided Video Self-Supervised Learning for Autonomous Driving

V-JEPA4A is introduced, a domain-specialized variant of V-JEPA for autonomous driving that is pre-trained on publicly available driving videos with a novel saliency-driven masking policy that preserves and predicts context according to semantic importance and temporal relevance, yielding more informative representation...

Christopher Lang, Alexander Braun, Abhinav Valada · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.