CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation
While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate explicit dynamics priors to improve physical plausibility, th...