Robo-Harness K1 is introduced, a robot-use agent (RUA) framework that exposes perception as tools that makes 3D geometry accessible without changing the VLM architecture or training a depth encoder, and suggests that perception-augmented RUAs offer a promising route to sample-efficient, generalizable robotic policies t...
Ze-Xi Li, Ye-Hang Zhang, Wen-Qian Li et al.· 0 citations
World Action Agent is presented, a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.
Ye-Hang Zhang, Hao-Jian Huang, Yi-Fan Chang et al.· 0 citations
Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams, and curate and structurally annotate 4.3 million physics images, carry expert-level annotation, and adapt a unified multimodal backbone.
Ming-Hui Zhang, Jinxin Shi, Yi-Fan Chang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.