Disentangle and Drop: Robust Universal Removal of Image Watermarks via Reconstructive Grayscale Residual Decomposition
Abstract
Invisible image watermarks are commonly evaluated against benign postprocessing operations such as compression, resizing, blur, and color changes. These tests leave out a different threat: a learned remover that preserves semantic image content while discarding residual evidence that carries the payload. We propose Disentangle and Drop (DnD), an attack that is agnostic to the watermark method and treats watermark removal as a representation routing problem. DnD decomposes a watermarked image into a semantic grayscale carrier and an auxiliary residual branch, and then suppresses the residual branch to reduce watermark evidence. The model is trained with latent spectral perturbations and low-strength diffusion exposure so that the drop operation remains stable under adaptive reconstruction. Experiments on seven representative watermark families show that one shared operating setting gives competitive removal with high visual fidelity. Operating scans and ablations separate usable attacks from image-damaging settings: stronger noise or diffusion can raise removal scores by damaging the image, while the practical regime comes from dropping the residual latent. These results argue for evaluating watermark robustness against learned removal at the representation level, not only against conventional image edits.