Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-distribution edge cases. Traditional active learning and manual annotation fail to scale against the complexity and volume of novel multimodal threats. In this paper, we propose an automated, agentic red-teaming framework that systematically synthesizes difficult examples using an iterative strategy that proposes novel hypotheses as well as mutating on past attempts. Leveraging a multi-agent architecture that consists of a high-reasoning Architect agent, an advanced image generator, and a multi-level verification committee of LLM raters, our system autonomously uncovers boundary-pushing violations and ambiguous policy edge cases without any human intervention. By employing these carefully synthesized adversarial examples as in-context demonstrations via test-time Retrieval, we substantially improve the target model's robustness, reducing the False Negative Rate (FNR) from 41.2% to 24.5% in a public image safety benchmark without relying on any human labeling.
This work introduces an evolutionary, task-agnostic, strategy-guided, executably-checkable data synthesis framework that, from minimal seed supervision, jointly synthesizes problems, diverse candidate solutions, and verification artifacts, and iteratively discovers strategies via a consistency-based evaluator that enforces agreement be-tween human-annotated and strategy-induced checks.
He Du, Bowen Li, Aijun Yang et al.· Annual Meeting of the Associ...· 0 citations
A failure-aware adversarial retrieval-augmented framework for improving robustness in natural language understanding that combines retrieval, automated validation, contextual-bandit failure selection, and controlled adversarial retraining, enables scalable robustness improvement without additional human annotation.
Roie Kazoom, Ofir Cohen, Rami Puzis et al.· 0 citations
This paper proposes a novel adversarial strategy, namely Prompt-Optimized Parameter Shaking (POPS), aiming to recover the supposedly unlearned multi-modality knowledge from the MLLMs, exposing fundamental vulnerabilities that challenge the foundational robustness of representative MMU-based privacy protections.
Zhangheng Li, Jianing Zhu, Junyuan Hong et al.· 0 citations
This position paper observes that a wide variety of techniques designed to improve specific aspects of LM behavior-targeting properties as diverse as adversarial robustness and factual coherence-can be understood as special cases of a common "consistency optimization" procedure and addressed with a standard set of optimization tools.
Itamar Hagay Pres, Belinda Z. Li, L. Ruis et al.· 1 citation
Agentic Data Evolution is proposed, a data-centric framework that organizes synthetic supervision as evolving data snapshots through a closed-loop Observation-Variation-Selection procedure, where a steady-state admission mechanism acts as a quality ratchet that conservatively gates updates for sustained cross-round improvement.
Yang Yu, Yilin Jiang, Zexuan Fei et al.· 0 citations