Large language model agents have demonstrated promising capabilities in cybersecurity tasks, yet their ability to reconstruct complete Advanced Persistent Threat attack campaigns from complex security logs remains largely unexplored. Existing cybersecurity benchmarks for agents mainly focus on vulnerability discovery,...
Qi Chen, Fu-Shuo Huo, Hang-Li Shen et al.· 0 citations
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this l...
Yuhao Pan, Haosong Peng, Zhengsheng Zhang et al.· 0 citations
The Unified Embodied Seeking and Following Benchmark (UESF-Bench) is introduced, a large-scale and diverse benchmark for embodied human seeking and following that requires agents to handle semantic-guided exploration, reliable behavior switching and recovery, and delayed identity grounding.
Kun Yu, Jianhua Yang, Yixiang Chen et al.· arXiv.org· 0 citations
This work proposes Progressively Disentangled and Recurrent Prompt Tuning (PDRPT), an edge-efficient framework that decouples object and state updates before joint refinement, suppresses traction force from highly-entangled prompts, and preserves alignment with the natural language space of CLIP.
Xiaocheng Lu, Chuan He, Ziming Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.