Robot foundation policies predict action chunks, but how many actions to execute before replanning depends on the current task phase. We introduce ChunkTrust, which treats the execution horizon as a latent variable inferred from action-expert evidence rather than a fixed hyperparameter. Its training-free Action-aware H...
Fan-Ding Huang, Jing-Yan Jiang, Shi-Feng Bao et al.· 0 citations
As large multimodal models move from understanding content to operating on digital environments, mobile GUI has emerged as a challenging and consequential testbed for digital embodied intelligence. Mobile agents operate under three coupled constraints: precise perception of complex interfaces, scalable acquisition of h...
Hy Vision Team, Huawen Shen, Zhengyang Tang et al.· arXiv.org· 1 citation
JarvisHub is introduced, a canvas-native creative agent harness for long-horizon multimodal creation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.