CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
CL4D is the first foundational 4D vision encoder that directly operates on dynamic point clouds, trained with a contrastive learning objective to align spatio-temporal geometric representations with natural language descriptions, and 4DVLM, a 4D vision-language model that conditions language generation on dynamic geome...