Skip to content

Author

Jiayu Ding

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models

As Vision-Language Models (VLMs) tackle dynamic 3D spatial reasoning, ego-motion perception becomes essential to resolve monocular scale ambiguity. However, current models often overfit to smooth trajectory priors rather than genuinely understanding physical motion. Consequently, their spatial reasoning degrades severely under large displacements, a phenomenon we term Kinematic Collapse. This failure stems from spurious visual-motion correlations in natural videos and a lack of explicit physical supervision. To evaluate this, we introduce Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties. Furthermore, we propose the TempoVista framework, featuring the Kinematic-GSPO algorithm. By embedding metric physical ground truth into policy optimization, TempoVista explicitly grounds visual representations in 3D space. Experiments demonstrate that our approach significantly improves both motion estimation and robust spatial reasoning by utilizing camera dynamics as an effective geometric calibration signal.

Jiayu Ding, Zhuo-Dong Liu, Lei Zhang et al. · 0 citations
Preprint Aug 2026

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting

While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle to interpret implicit intents, complex spatial constraints, and commonsense reasoning required for practical embodied interactions. To address this gap, we introduce the task of reasoning 3D Gaussian segmentation and construct two benchmarks, Causal-LERF and Causal-ScanNet. These benchmarks systematically evaluate commonsense, spatial, affordance, and counterfactual reasoning. Evaluations reveal that current state of the art methods perform poorly on these reasoning challenges. Therefore, we propose CausalSplat, a framework that integrates vision-language models with 3D scene graphs to disentangle explicit structural perception from implicit logical inference. Extensive experiments demonstrate that CausalSplat achieves state of the art performance on our reasoning benchmarks while showing strong generalizability on standard referring and open vocabulary 3D segmentation tasks. Project Page: https://jiayuding031020.github.io/CausalSplat

Jiayu Ding, Meilu Song, Yun Chen et al. · 0 citations
Preprint Jul 2026

ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting

ZeroSplat lifts 2D Vision-Language Model priors into 3D space through robust multi-view geometric constraints and enables intrinsic point-level understanding without incurring any additional feature storage, and significantly outperforms state-of-the-art methods across generalized and single-target scenarios while maintaining exceptional efficiency.

Jiayu Ding, Meilu Song, Xiaoyi Zhang et al. · 1 citation