Less Language, More Latents: Annotation-Efficient VLAs for Driving
A three-stage pipeline that transforms abundant unlabelled observation-trajectory pairs into a substrate for language-conditioned control and achieves a Driving Score of 87.98 and a Success Rate of 70.46%, matching or surpassing fully supervised baselines on the closed-loop Bench2Drive benchmark.