While Vision-Language-Action (VLA) models demonstrate impressive capabilities in robotic manipulation, their memoryless nature renders them brittle to test-time environment shifts, particularly hardware shifts caused by wear or imperfect calibration. Enabling these models to self-adapt during deployment without requiri...
Hong-Xin Zhang, Chun-Tse Lin, T. Wang et al.· 0 citations
Text2Sim is presented, a simulation-specialized agentic pipeline that converts a text-only request into an executable, editable dynamic case and achieves higher scores than all four state-of-the-art baselines on both metrics.
Xiao-Yu Xiong, T. Wang, Yi-Ling Qiao et al.· 0 citations
GS-Agent is presented, an end-to-end multi-agent framework that integrates physics engines in the loop to generate realistic, dynamic, and controllable 4D physical worlds from natural language, envisioned as a foundation for a new paradigm in 4D world generation, empowering creative content creation and physical AI.
Hongxin Zhang, Chun-Tse Lin, Junyan Li et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.