Skip to content
Preprint

MarkSec: Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

Sep 2026 · 0 citations · 37 references
Computer Science

TL;DR

Results show that apparent attack winners depend on text-quality constraints, attack generality, and capability assumptions, and that apparent attack winners depend on text-quality constraints, attack generality, and capability assumptions.

Abstract

LLM watermarking helps trace the origin of generated text, but faces stealing attacks that recover watermark information, scrubbing attacks that remove watermark signals, and spoofing attacks that forge text accepted as watermarked. These attacks are often studied in isolation, leaving their connections unclear. Evaluations also often lack shared detector calibration, metric definitions, and reporting protocols. Moreover, measuring attack success and text quality separately makes it difficult to identify attacks that are both effective and quality-preserving. We propose MarkSec, a general framework that unifies analyses of stealing, scrubbing, and spoofing. We evaluate attacks under a common reporting protocol and introduce a quality-constrained attack success metric to assess effectiveness and text quality jointly. Experiments across representative watermark families, attacks, LLMs, and datasets reveal three findings. First, attacks that appear strongest by watermark removal alone can fall behind general rewriting when success also requires acceptable text quality. Second, general rewriting remains a strong baseline across watermark families, while its advantage over other scrubbers varies by family. Third, in a case study of one watermark family, stealing-based scrubbers often underperform the best general-scrubbing baselines when text quality is required. These results show that apparent attack winners depend on text-quality constraints, attack generality, and capability assumptions.

View source

Similar papers

Preprint Sep 2026

RemTraceNet: Few-Shot Forensic Detection of Invisible Watermark Attacks

Removing an invisible watermark and concealing the forensic evidence are distinct objectives: successfully disrupting the embedded watermark does not imply that the removal process is forensically undetectable. When verification fails, removal traces can provide complementary evidence for provenance and ownership verif...

Ji-Dong Yang, Hu Yu, Qi Li et al. · 0 citations
Preprint Sep 2026

DeMark: A Query-Free Black-Box Attack for Quality-Preserving Audio Watermark Removal

Audio watermarking protects digital speech by embedding imperceptible signals for ownership verification and misuse tracing. However, the security of learning-based watermarking remains insufficiently understood under realistic adversarial removal, where attackers cannot access or query the watermark encoder, decoder,...

Wei-Kang Ding, Bin-Hao Ma, Han-Qing Guo et al. · 0 citations
Preprint Sep 2026

FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models

Text-to-image diffusion models enable data-efficient"mimicry"attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are br...

Hao-Yang Li, Ruo-Xi Sun, Qing-Qing Ye et al. · 0 citations
Preprint Aug 2026

How Fragile Is Your Watermark? Training-Free Structural Removal of Neural Audio Watermarks

Neural audio watermarks are increasingly used to attribute and detect AI-generated speech, so their practical value rests on how cheaply an adversary can remove them. Robustness is usually measured by running a fixed battery of distortions blindly against every scheme. We instead make removal diagnostic: from a few cle...

Likhith Kumara · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.