Skip to content
Preprint

SPARK: Representation-Level KV Memory Alignment for Safer Vision-Language Models

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

SPARK is introduced, a two-stage framework for targeted KV-memory repair that reduces multimodal attack success while preserving general capability and suggests that multimodal jailbreak behavior can be substantially mitigated by selectively repairing intervention-relevant KV subspaces at prefill.

Abstract

Vision-language models (VLMs) remain vulnerable to jailbreaks that distribute harmful intent across text and images, making unimodal safety mechanisms insufficient. We investigate whether this vulnerability can be mitigated directly in the multimodal key-value (KV) memory formed during prefill, without modifying model parameters at inference time. We introduce SPARK, a two-stage framework for targeted KV-memory repair. Stage 1 uses a disposable diagnostic adapter to identify harm-associated directions in multimodal key and value representations. Stage 2 projects out these directions, learns a lightweight residual repair, and anchors repaired keys with an image-structural prior to preserve visual grounding. Rather than applying the intervention uniformly, SPARK mixes repaired and original memory using a head-wise coefficient g_h* determined by intervention-relevant subspace energy E_h, requiring no explicit harm classifier at inference. Across LLaVA-OneVision-7B, Chameleon-7B, Qwen2-VL-7B, and InternVL2-4B, SPARK reduces multimodal attack success while preserving general capability. On LLaVA-OneVision-7B, image-only jailbreak attack success falls to 4.7%, while MMMU remains within 0.6 points of the undefended model (47.8 vs. 48.4) with near-baseline language quality. On MM-SafetyBench, attack success decreases from 39.2% to 12.4%. Even under white-box adaptive joint prompt-image attacks, attack success is limited to 20.3%, compared with 54.6% for the undefended model. These results suggest that multimodal jailbreak behavior can be substantially mitigated by selectively repairing intervention-relevant KV subspaces at prefill, particularly when harmful evidence is carried by the visual modality.

View source

Similar papers

Preprint Aug 2026

COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models

It is suggested that multimodal safety cannot be enforced reliably without modeling the requested operation, the visual target to which it applies, and the confidence of that grounding.

Md Abdullahil Oaphy, Anhao Xiang, Zongxing Xie et al. · 0 citations
#machine learning Preprint Aug 2026

Do VLMs Share Safety Neurons Across Modalities?

A causal, neuron-level analysis of safety mechanisms in 10 VLMs, a two-stage detection pipeline with iterative ablation that accounts for self-repair, and two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval, which decouple visual and textual safety signals are introduced.

Jia-Xuan Li, Jia-Hao Zhang, D. Vo et al. · 0 citations
#machine learning Preprint Sep 2026

Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis

It is found that post-pretraining QK-RMSNorm injection fails to reproduce the protection of native QK-RMSNorm, while several off-the-shelf weight-merging settings fail to recover the lost capability after VL training, highlighting the value of screening backbones with Sink Strength before VL training and narrow the int...

Minsik Choi, Geewook Kim, Young Geun Kim · 0 citations
Preprint Aug 2026

VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference

VisCache is proposed, a plug-and-play framework for coarse-to-fine KV pruning without training that consistently outperforms existing baselines, establishing a new Pareto frontier between efficiency and performance for long-context VLLM inference.

Lyuke Wang, Zhuo Li, Guang-Xu Zhu · 0 citations
Preprint Aug 2026

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capabil...

Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.