Evaluating language-guided mobile agents has recently shifted from rule-based to model-based approaches to achieve scalable and automated assessments. However, existing holistic evaluation paradigms process entire trajectories at once, leading to substantial context overload. Moreover, they primarily focus on task comp...
Peng-Jian Yang, Zijing Gao, Xue Yu et al.· 1 citation
Agent memory allows LLM agents to use earlier interactions when answering new queries. Existing methods often compress interaction histories into summaries or other LLM-generated representations. Repeated generation adds cost and can discard answer-bearing details before the system knows what a future query will requir...
Ze-Yang Cui, Jian-Nong Cao, Zhiyuan Wen et al.· 0 citations
InstructVVT is proposed, an instruction-driven and reference-guided video virtual try-on framework based on a Diffusion Transformer that operates without inference-time spatial priors that outperforms state-of-the-art open-source methods in garment fidelity, structural preservation, and temporal consistency, despite re...
Di Shao, Song-Han Wu, Xin-Yu Chen et al.· 0 citations
This paper introduces SeerGuard, a consequence-aware safety framework designed to mitigate risks through pre-execution instruction-level screening and action-level risk assessment, and constructs a unified safety-augmented world model (SAWM) via multi-task learning, integrating semantic next-state prediction with safet...
Xue Yu, Bo Yuan, Peng-Jian Yang et al.· arXiv.org· 1 citation
World Tokens is an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeling, and action generation and is highly competitive on LIBERO, attains the best reported averages on SIMPLER, and substantially improves real-world R1 Pro success over a matched...
Qu Tang, Benhui Zhuang, Bo Yuan et al.· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.