Qualitative examples show the PhysBrain 1.5 model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.
DeepCybo Team, Yue Bin, Hai-Peng Cao et al.· 0 citations
Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet sma...
Shijie Lian, Bin Yu, Zhaolong Shen et al.· 1 citation
IntentVLA is introduced, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation and improves rollout stability and outperforms strong VLA baselines.
Shijie Lian, Bin Yu, Xiaopeng Lin et al.· arXiv.org· 6 citations
SIEVE, a structure-aware data selection method for VLA imitation learning that can surpass full-data training while using only 50% of demonstrations and 50% of training steps, suggests that reusable structure, captured through primitives and transitions, is an important signal for efficient VLA imitation learning.
Changti Wu, Bin Yu, Zhaolong Shen et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.