Recent vision-language-action (VLA) policies demonstrate promising generalization across diverse short-horizon tasks. However, they remain unreliable on long-horizon tasks, partly because the large-scale training data is biased toward single-stage manipulation tasks that are cheaper to demonstrate. A single weak atomic...
Zi-Wen Li, Hanlue Zhang, Zhen-Yang Ren et al.· 0 citations
Autoregressive pretraining increasingly draws on heterogeneous data, making it important to understand how a model learns from an individual token. The next-token prediction objective naturally identifies a token's contribution with its own loss. However, each token is not only a prediction target but also context for...
Su-Qin Yuan, Runqi Lin, Ke-Yu Lin et al.· 0 citations
It is shown that preference alignment preserves the human response distribution only under a restrictive condition, and no consistent evidence that real human preferences satisfy it, and human-likeness is established as an explicit dimension of alignment rather than something assumed to follow from preference alignment...
Su-Qin Yuan, Runqi Lin, Mu-Yang Li et al.· 0 citations
Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a distribution mismatch with production environments; expensive real-device failures remain underutilized; and fixed benchmarks saturate, losi...
Tong Ye, Kunyang Han, Guo-Zhi Wang et al.· 0 citations
GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), is introduced, revealing the substantial gap between current agent capabilities and those required for complex...
Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.