Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampl...
Jesson Wang, Shawn Li, Wei Yang et al.· 0 citations
Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, lea...
P. Wang, Chen-Hao Liang, Ze-Long Xu et al.· 0 citations
Large language model (LLM) agents increasingly rely on external skills, but routing user requests over large skill registries is difficult because many skills are functionally redundant while complex tasks often require complementary skill sets. Existing skill routers typically rank candidates independently by query re...
Wang Wei, Tiankai Yang, Samyadeep Basu et al.· 2 citations
AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.
Han Bao, Yue Huang, Yan-Bo Wang et al.· Proceedings of the 32nd ACM...· 0 citations
SocialMaze is introduced, a benchmark that organizes six tasks across social deduction games, daily-life interactions, and digital community platforms along three descriptive design axes: deep reasoning, dynamic interaction, and information uncertainty.
Zi-Xiang Xu, Yan-Bo Wang, Yue Huang et al.· 1 citation
A benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth is introduced, establishing scalable geometric reasoning as an open challenge for vision-language models.
DOG-DPO is proposed, a training-free data selection framework that treats preference pairs as structured geometric signals and recovers most of the safety gains of full-data training while remaining entirely teacher-free, training-free, and substantially faster than representative selection baselines.
Yi Nian, Tiankai Yang, Yudi Zhang et al.· arXiv.org· 1 citation
WeClawArena is introduced, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.
P. Wang, Ao-Jie Yuan, Haiyu Zhang et al.· 1 citation
CatchBench puts one auditor's question to three information states: the declared configuration before a run (PRE), a growing prefix of its trace (LIVE), and the finished trace (POST), which none scores all three under one task-method interface.
Yue Zhao, Meng-Yuan Li, Ruo-Lin Li et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.