Model-based control can directly execute specified objectives, while learning can amortize such behaviors into reactive policies, making their combination a natural solution to multi-stage manipulation. We introduce Semantically UNified (SUN) Programs, typed executables that compile grounded relations into aligned opti...
Wei-Qi Wang, Zhi Li, Yuliang Lei et al.· 0 citations
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe...
Wei-Qi Wang, Yu-Xin Zhou, Mou-Xiang Chen et al.· 0 citations
The results establish that simulation-screened task semantics can effectively amortize control into robust policies, without demonstrations or manual dense rewards, unifying symbolic planning and data-driven execution.
Wei-Qi Wang, Zhi Li, Yuliang Lei et al.· 0 citations
Experiments show that UI-Mate-27B sets a new open-weight state of the art on general computer-use benchmarks, substantially improving long-horizon reliability, and makes three contributions to an environment-grounded training stack with in-context demonstration learning.
Zihan Ding, Longxu Dou, Qixiao Gao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.