Skip to content
Preprint

CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

Aug 2026 · 0 citations · 46 references
Computer Science

TL;DR

CL4D is the first foundational 4D vision encoder that directly operates on dynamic point clouds, trained with a contrastive learning objective to align spatio-temporal geometric representations with natural language descriptions, and 4DVLM, a 4D vision-language model that conditions language generation on dynamic geometric representations.

Abstract

4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vision encoders are largely limited to static 2D images or 3D point clouds without temporal modeling, or to 2D videos that lack accurate geometric depth reasoning. Consequently, current approaches fail to jointly capture spatial structure and motion evolution in dynamic scenes. We present CL4D, the first foundational 4D vision encoder that directly operates on dynamic point clouds, trained with a contrastive learning objective to align spatio-temporal geometric representations with natural language descriptions. By learning a shared embedding space between text and 4D scene dynamics, CL4D enables zero-shot motion-to-text and text-to-motion retrieval in dynamic environments and serves as a foundational 4D vision encoder for downstream 4D vision-language tasks. Building on this encoder, we introduce 4DVLM, a 4D vision-language model that conditions language generation on dynamic geometric representations. 4DVLM is the first VLM designed to operate directly on 4D point clouds without relying on 2D images, 2D videos, or static 3D point clouds. We train CL4D and subsequently 4DVLM on a newly constructed dataset termed DynAction4D capturing diverse human motions across varying object interactions and scene environments. Extensive experiments across multiple 4D human action benchmarks demonstrate that CL4D achieves state-of-the-art performance, with improvements of approximately ~16.75% over prior methods. Furthermore, 4DVLM outperforms frontier video VLMs such as Gemini and GPT-5 even when these models are provided with RGB video sequences corresponding to the same scenes represented as 4D point clouds for 4DVLM.

View source

Similar papers

Preprint Aug 2026

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Qwen-3D is introduced, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes and incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene repre...

Lucy Lin, Ayush Jain, Yifan Liu et al. · 3 citations
Preprint Aug 2026

StrataVLA: Hierarchical and Efficient 3D Geometric Grounding for Vision-Language-Action Models

StrataVLA is introduced, a plug-and-play framework for hierarchical geometry grounding that achieves 98.53% average success on LIBERO suites while reducing geometry-model invocations by up to 88%, establishing hierarchical geometry injection as an effective and efficient way to achieve spatially grounded robotic contro...

Jin Cui, Zhao Pu, Bo Cai et al. · 0 citations
Preprint Sep 2026

GeomVLA: Unifying Scene, Motion, and Action in 3D

We present GeomVLA, a Vision-Language-Action (VLA) model that unifies perception, latent scene motion prediction, and action generation within a shared robot-centric 3D coordinate frame. Our approach lifts pretrained VLM features into spatially grounded 3D scene tokens using depth and camera calibration, while retainin...

Zi-Yin Xiong, Nikolaos Gkanatsios, Moritz Reuss et al. · 0 citations
Sep 2026

EgoMotion: Hierarchical Vision-Language Learning and Diffusion for Egocentric Motion Generation.

EgoMotion is proposed, a two-stage framework for vision-language-guided egocentric motion generation that achieves state-of-the-art performance and produces motion sequences that are both semantically grounded and kinematically superior to existing approaches.

Rui-Bing Hou, Ming-Yu Zhou, Yu-Wei Gui et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models

Vision-language-action (VLA) models built on pretrained vision-language models have demonstrated strong performance across diverse robotic manipulation tasks. However, VLA models that directly map current 2D observations to actions often lack sufficient spatial and temporal understanding, limiting their performance in...

Zhi-Yuan Gao, Di Wen, Yan Zhan et al. · 0 citations
Preprint Sep 2026

Geometric Encoding for Spatial Reasoning in Vision-Language Models

Vision-Language Models (VLMs) are far more reliable at recognizing what appears in a video than at reasoning about its spatial and temporal properties, such as metric distances, object dimensions, and consistent object identities across frames. We present Geometric Code, a perception-to-geometry pipeline that computes...

Antonio Jun, Hao-Shui Yu, Zheng Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.