SIRModel: Learning Spatial Intermediate Representation to Parameter-Efficiently Fine-Tune a Vision Language Model for Manipulation
Long-horizon robotic manipulation requires a policy to bridge task-level semantic reasoning with metric three-dimensional interaction geometry. Existing vision–language–action policies usually acquire geometry implicitly from visual tokens or introduce deterministic intermediate variables only in the image plane, which...