This work proposes MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios, harm severity, and modality-specific stealth levels and introduces a hierarchical evaluation framework to assess fundamental response reliability, actual risk exposure, and the structural integrity of defensive behaviors.
Abstract
While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related annotations and rely on unidimensional metrics, hindering comprehensive robustness evaluation. To address this, we propose MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios, harm severity, and modality-specific stealth levels. Furthermore, we introduce a hierarchical evaluation framework to assess fundamental response reliability, actual risk exposure, and the structural integrity of defensive behaviors. Extensive zero-shot evaluations across 17 state-of-the-art MLLMs provide a comprehensive safety profile of current multimodal systems. Our analysis systematically investigates cross-modal input configurations and uncovers safety implications associated with Chain-of-Thought (CoT) reasoning. These multifaceted findings underscore the urgent need for robust, reasoning-aware safety alignment in the multimodal landscape.
MMJailBench is introduced, a factorized benchmark that systematically varies and combines these factors under controlled configurations, enabling fine-grained comparison and factor-level attribution and a modular multimodal jailbreak evaluation suite with full and lightweight configurations, multiple judge options, and...
Tian-Shi Wang, Jing-Song Wang, Ya-Fei Huang et al.· 0 citations
This work proposes an explanation-aware safety framework that augments binary harmfulness detection with structured, human-interpretable explanations capturing severity, strategies, trigger spans, ratio-nales, and derived safety factors, and introduces a human–LLM hybrid annotation and canonicaliza-tion pipeline.
Sunghee Dong, Sungwon Yi, K. Bae et al.· Proceedings of the Thirty-Fi...· 0 citations
ReFrame is a training-free multimodal input reframing framework where two agents share a lightweight locally deployed MLLM: the evidence-generation agent constructs complementary risk and utility evidence, and the rewrite-and-routing agent converts it into a safe proxy prompt and image-routing decision before calling t...
Wenzheng Jiang, Xuan-Kun Rong, Yuan-Zhao Zhai et al.· 0 citations
A mechanistic analysis of the jailbreaking behavior in a large-scale, safety-aligned LLM, focusing on LLaMA-2-7B-chat-hf identifies computational circuits responsible for generating affirmative responses to jailbreak prompts and uncovers key attention heads and MLP pathways that mediate adversarial prompt exploitation.
Paria Mehrbod, Boris Knyazev, Guy Wolf et al.· 0 citations
Large language models (LLMs) have demonstrated remarkable capabilities across languages, yet their safety
alignment remains predominantly evaluated in monolingual, especially English, settings. Code-switching (also known as codemixing)—the alternation between two or more languages within a single utterance or conversat...
Pulagam Naveen Kumar, Sowjanya Bojja, K. Anoosha et al.· International Journal for Re...· 0 citations
EADC, a novel advanced evaluation benchmark of LLMs based on an AI compliance knowledge graph and AI compliance legal experts, is introduced, offering a rigorous, AI laws and regulations-aligned benchmark to safeguard high-level and deep compliance in the application of LLMs.
Yan Zhang, Rui-En Li, Yao-Yao Peng et al.· 0 citations
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.