Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study adversarial images that preserve full-token correctness yet induce errors after compression, even when both inference paths succeed on the clean image. Creating such fail...
Shilinlu Yan, Bo-Wen Chen, Yuechen Zhang et al.· 0 citations
Large reasoning models (LRMs) improve performance on complex tasks through extended reasoning, yet the same process can degenerate into redundant verification and persistent generation loops. Such uncontrolled reasoning increases inference cost and creates risks of resource exhaustion and service degradation. However,...
Yuan-He Zhang, Ziwei Wang, Jie Ren et al.· 0 citations
Results show that two-stage defense feedback can serve as a useful learning signal for adaptive red teaming and that evaluating either defense stage in isolation can miss the resulting attack capability.
Peng-Yu Zhu, Jing-Yi Yang, Yi Liu et al.· 0 citations
This work presents UniACE, a unified framework for model-centric evaluation under an explicit, common execution condition, and reports agent benchmark outcomes as properties of an explicit evaluation configuration, enabling more interpretable and reproducible cross-benchmark comparisons.
Peng-Yu Zhu, Lijun Li, Yaxing Lyu et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.