The high visual-token load in multimodal large language models (MLLMs) motivates training-free pruning to reduce later-layer computation, but under a fixed budget, pruning must preserve query-relevant evidence while avoiding redundancy. Existing methods rank tokens, diversify selected subsets, or optimize coverage with...
Sheng-Li He, Yong-Chao Liang, Rou-Meng He et al.· 0 citations
CAT is introduced, a paradigm in which the agent writes Playwright code to drive the browser, gathers feedback, and autonomously explores web applications to uncover bugs, revealing a clear gap between current VLM capabilities and the demands of real-world testing in AI web development.
Bin Hong, Zhen-Chao Zhang, Ji-Yuan He et al.· 0 citations
Meta-Task is proposed, a framework that redefines terminal task synthesis as a Terminal-Bench-format task itself: an agent operates within a real container environment to iteratively generate, execute, and verify tasks, so that synthesized components are checked for internal consistency and executability within the gen...
Zhi-Hong Pan, Ji-Yuan He, Kai Zhang et al.· arXiv.org· 1 citation
ICAE-Bench, a benchmark for evaluating coding agents under interactive project-building settings, starts from a fuzzy product requirement, simulating the dynamic paradigm with an automated User Agent, and introduces three key designs.
Zhongyuan Peng, Dan Huang, Chuyu Zhang et al.· arXiv.org· 3 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.