Reliable evaluation of image forgery localization (IFL) requires assessing models under diverse distribution changes, yet existing benchmarks often cover limited manipulation conditions or entangle multiple factors in cross-dataset evaluation. Consequently, aggregate performance provides an incomplete view of localizat...
Bao-Ke Dou, Zi-Ye Wang, Hao Wang et al.· 0 citations
LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition, and relationships. Agents must complete legitimate tasks and prevent unauthorized disclosure of p...
Hao Chen, Wen-Hui Dong, Ye Chen et al.· 0 citations
This work proposes MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios, harm severity, and modality-specific stealth levels and introduces a hierarchical evaluation framework to assess fundamental response reliability, actual risk exposure, a...
Yi Shi, Yueming Lyu, Hao-Xiang Tan et al.· 0 citations
This work introduces PRMU, a benchmark for evaluating corpus-free multimodal unlearning under realistic person-centric deletion requests, and introduces Similarity-Gated Projection Editing (SGPE), a lightweight corpus-free unlearning baseline with knowledge displacement, protected parameter-space editing, and locality-...
Hua-Feng Chen, Yueming Lyu, Ziyuan Chen et al.· 0 citations
AURORA-LM is introduced, a continuous-latent diffusion language model that separates the construction of a decodable text representation from the modeling of its distribution, and achieves the strongest performance among evaluated continuous and diffusion-based language models on OpenWebText free generation and XSum su...
Jiajun Liang, Yu-Ling Liao, Yu-Kang Cao et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.