Experiments demonstrate that DDPO significantly outperforms static prompt optimization methods, particularly on weakly aligned models and when handling semantically ambiguous benign prompts, successfully distinguishing them from genuinely harmful requests.
Doniyorkhon Obidov, H. Yu, Xiaolong Guo et al.· AAAI Conference on Artificia...· 4 citations
It is demonstrated that an attacker can embed a stealthy backdoor into an LLM-based robot controller by manipulating its instructions, triggered not by an external cue, but by a specific, rare sequence of the robot's own past actions.
Doniyorkhon Obidov, Shivayogi Akki, Cheng-Qiu Tan et al.· 2 citations
CAP is introduced, a scalable benchmark for evaluating browser agents on cross-site, human-like web tasks that require non-trivial UI interactions and visual understanding and a decomposition-and-recomposition pipeline that first abstracts each website into a structured site card capturing user-facing functions, comple...
Zejun Xu, Taiyi Chen, Jin Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.