Resolving State-Representation Mismatch: State-Space Visual Reasoning for Open-Loop VLA Planning
Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pixel-, and latent-spac...