Skip to content

Author

Shenggang Wei

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#diffusion models Open access Aug 2026

GSP-3D: generalizable 3D Gaussian Splatting with diffusion policy and world model for long-horizon bimanual manipulation

Introduction Learning robust visuomotor policies for bimanual manipulation remains challenging due to the stringent requirements for precise coordination between arms and the ability to generalize across diverse environmental conditions. Existing diffusion-based policies often suffer from temporally inconsistent action generation, while their reliance on sparse point cloud representations limits structural completeness and fails to capture future scene dynamics, hindering performance in long-horizon tasks. Methods To address these limitations, we introduce GSP-3D, a unified framework that integrates generalizable 3D Gaussian Splatting (3DGS) with diffusion-based policy learning. GSP-3D comprises three key components: (1) a Generalizable Gaussian Regressor (GGR) that predicts 3D Gaussian primitives from a single RGB-D frame in real time; (2) a 3DGS-conditioned diffusion policy that aggregates Gaussians into compact latent representations, replacing point clouds with explicit geometric primitives; and (3) a transformer-based world model that forecasts future Gaussian sets and leverages prediction error as an auxiliary loss to enforce temporal consistency across actions. Results We evaluate GSP-3D on the RoboTwin 2.0 benchmark across a range of bimanual manipulation tasks under both clean and domain-randomized conditions, including variations in lighting, backgrounds, and tabletop distractors. Experimental results show that GSP-3D consistently outperforms existing baseline methods, achieving higher success rates while maintaining minimal computational overhead. Discussion These findings demonstrate that integrating explicit 3D Gaussian representations with diffusion policies offers an efficient and robust solution for long-horizon, temporally coherent bimanual manipulation, effectively addressing the generalization and consistency challenges that limit current approaches.

Xukun Liu, Junhua Huang, Shenggang Wei et al. · 0 citations