While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visua...
Yan-Yan Zhang, Di-Sheng Liu, Xin-Peng Li et al.· 0 citations
Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and...
Shouren Wang, Chuan Ma, Mohsen Hariri et al.· 0 citations
It is shown that a lightweight LLM-only DPO update on tiny single-object-pair synthetic data mitigates the bias, lifting four-way robust accuracy by up to 100 points on synthetic data, and by 68.1 points on broader evaluation datasets WhatsUp, SpatialMQA-Direct, and VSR.