VPTwin is proposed, a Real-Sim-Real video prediction framework that anchors real-world future prediction using real-synchronized simulation twins and establishes a predictive planning loop using VPTwin to visually verify VLM-proposed actions and guide reliable real-world execution.
Zheng-Hao Xiao, Min-Ting Pan, Nan-Tian He et al.· 0 citations
Biological wet-lab protocols are written for trained researchers and often leave routine operations, state-dependent conditions, and contextual parameters implicit, making them difficult to translate into robot-executable actions. We present ProtoAct, a structured protocol-grounding framework that converts free-form bi...
Zhe Liu, Jiaming Gu, Zhao-Hui Du et al.· 1 citation
This work introduces BioVLN, a simulation platform for developing and evaluating visual-language navigation agents in biomedical laboratories and shows that geometric exploration reaches 74.4--87.5% success, while sampling multiple valid positions in the operation area improves success and reduces unsafe proximity.
Zhe Liu, Quan Lu, Zhao-Hui Du et al.· arXiv.org· 0 citations
LabRobFail, a failure-centric framework for learning and evaluating robotic failure analysis in chemical laboratories, and LabRobFail-VLM, a domain-specialized vision-language model that generates structured failure diagnoses and recovery instructions, demonstrate the value of fine-grained failure understanding for clo...
Haobo Wang, Baoli Sun, Anqi Zou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.