Erasing individual identities from Vision-Language Models (VLMs) is uniquely challenging because personal data is entangled across modalities rather than stored as isolated attributes. However, existing multimodal unlearning benchmarks primarily evaluate attribute-centric forgetting, overlooking the more critical objec...
Xiong-Tao Sun, Hui Li, Tian-Tong Wu et al.· 0 citations
Advancing beyond traditional static scoring models, LLM-powered agentic recommender systems (LLM-ARS) instantiate users and items as autonomous agents, whose semantic states are dynamically refined through a recurrent process known as collaborative reflection. While this mechanism improves recommendation quality, it si...
Yu-Rong Hao, Wen Zhou, Guo-Wei Guan et al.· 0 citations
FraudBench is a multimodal benchmark for detecting AI-generated fraudulent refund evidence and shows that current MLLMs often recognize real-damaged evidence but fail on many fake-damaged subsets, with fake-damage detection rates far below the 50\% baseline on most generator subsets.
Most studies of prompt injection focus on generative agents, leaving their effects on models with schema-defined outputs unclear. We examine these effects in Jev, a non-generative decision model, using 510 reconstructed InjecAgent cases. Malicious content shifts action probabilities but rarely causes Jev to select the...
REFLEX, an agent architecture that uses Jev as a fast, typed decision layer and calls a strong LLM when confidence is low, or generation is required, is studied, identifying when selective control with Jev can reduce computation and where its benefits are limited.
Results indicate that current LLMs are substantially more reliable at detecting broad collaboration opportunities than at identifying their specific types and role directions.
Tian Du, Tian-Tong Wu, Ya-Fei Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.