GeoLAM: Learning Geometry-Grounded Latent Actions from Unlabeled Human Videos
GeoLAM, a framework for learning geometry-grounded latent actions from action-free human videos, combines future-frame reconstruction through a frozen geometric feature hierarchy with motion supervision from a training-only 4D geometry teacher and requires neither the geometry teacher nor future-video generation.