Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score h...
Chang Guo, Yu-Kun Xie, Bo-Han Tan et al.· 0 citations
Long-context LLM training suffers from a load-balancing problem that sequence packing does not solve. Packing samples into fixed-token sequences balances memory and linear-cost operators, but the dominant attention cost scales with the sum of squared sequence lengths. Thus, equally sized packed sequences drawn from a l...
Yan Wang, Xiu-Long Yuan, Kaiming Yang et al.· arXiv.org· 1 citation
This work demonstrates that robust semantic grounding can be achieved through elegant structural design, bypassing the inefficient brute-force data scaling paradigm and introduces ParaVLA, a natively decoupled 0.33B-parameter model exhibiting near-perfect robustness to instruction rewording.
Zhao-Kai Yin, Zhi-Peng Zhang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.