KeyBound is presented, a learned audio watermark that restores the two ingredients classical watermarking supplied and learned schemes set aside, a secret key and a host-aware carrier, so the key governs payload access while the host-conditioned carrier resists direct transplantation.
Abstract
Audio watermarking is a proactive route to attributing synthetic speech to its source. Learned audio watermarks are typically judged by payload recovery after a fixed catalog of signal distortions such as noise, compression, filtering, and resampling. That test is necessary but not sufficient for provenance. A mark offered as evidence of origin should not be readable by an unauthorized party, should not be transferable to unrelated audio, and should not vanish when the recording is re-synthesized by a modern generative model. We present KeyBound, a learned audio watermark that restores the two ingredients classical watermarking supplied and learned schemes set aside, a secret key and a host-aware carrier. KeyBound masks the payload with a secret key and embeds the masked bits through a carrier modulated by a frozen spectral representation of the host, so the key governs payload access while the host-conditioned carrier resists direct transplantation. A key-independent presence head lets any party detect a mark, whereas only a key holder reads its attribution, and under the single-sample uniformity assumption a wrong-key decode clears our verification rule with probability at most $2.1\times10^{-3}$. On LibriSpeech against WavMark, AudioSeal, and Timbre, KeyBound holds 1.00 detection accuracy and 0.98 bit accuracy under a spectral denoiser that costs every baseline its detection, decodes at chance without the key, and rejects transplanted carriers. Detection further transfers to held-out DAC and BigVGAN re-synthesis, though exact payload recovery degrades. Speech provenance is thus better posed as a keyed, host-bound attribution problem than as the recovery of a payload under a catalog of signal distortions fixed in advance.
Audio watermarking protects digital speech by embedding imperceptible signals for ownership verification and misuse tracing. However, the security of learning-based watermarking remains insufficiently understood under realistic adversarial removal, where attackers cannot access or query the watermark encoder, decoder,...
Wei-Kang Ding, Bin-Hao Ma, Han-Qing Guo et al.· 0 citations
HaloMark, a watermark for embedding vectors cryptographically bound to a C2PA manifest, is presented, a watermark for embedding vectors cryptographically bound to a C2PA manifest bound rigorously for linear and non-adaptive attackers and characterise empirically for the adaptive case.
SpreadMark keeps the embedded watermark imperceptible, maintaining high perceptual quality on both COCO and DIV2K, and is the only evaluated method retaining high detection under both the regeneration and the latent-space sparsification settings the authors test.
Wei Song, Yu-Xin Cao, Zhen-Chang Xing et al.· 0 citations
Large language model (LLM) watermarking provides an important mechanism for tracing the provenance of generated text. Existing statistical watermarks are often effective and robust, but most of them rely on private detection keys, which centralizes verification and complicates public auditing. Recent public or publicly...
Haocheng Fu, Yuqi Qian, Luyao Wang et al.· 0 citations
Existing video watermarking systems are symmetric: the party that can verify a mark holds the extractor weights or generator secret and can therefore also embed one. Benchmarks confirm the consequence, reporting that white-box forgery defeats all evaluated methods. We present a training-free video watermark that remove...
Redwing is proposed, a design principle for robust token-level watermarking that generalizes to TTS models at a speech-quality cost close to that of KGW and shows that retokenization is not merely a source of noise: its transition structure can be exploited as a design principle for robust token-level watermarking.