Skip to content
Preprint

KeyBound: Keyed and Host-Bound Learned Audio Watermarking for Speech Provenance

Aug 2026 · 0 citations · 27 references
Computer Science

TL;DR

KeyBound is presented, a learned audio watermark that restores the two ingredients classical watermarking supplied and learned schemes set aside, a secret key and a host-aware carrier, so the key governs payload access while the host-conditioned carrier resists direct transplantation.

Abstract

Audio watermarking is a proactive route to attributing synthetic speech to its source. Learned audio watermarks are typically judged by payload recovery after a fixed catalog of signal distortions such as noise, compression, filtering, and resampling. That test is necessary but not sufficient for provenance. A mark offered as evidence of origin should not be readable by an unauthorized party, should not be transferable to unrelated audio, and should not vanish when the recording is re-synthesized by a modern generative model. We present KeyBound, a learned audio watermark that restores the two ingredients classical watermarking supplied and learned schemes set aside, a secret key and a host-aware carrier. KeyBound masks the payload with a secret key and embeds the masked bits through a carrier modulated by a frozen spectral representation of the host, so the key governs payload access while the host-conditioned carrier resists direct transplantation. A key-independent presence head lets any party detect a mark, whereas only a key holder reads its attribution, and under the single-sample uniformity assumption a wrong-key decode clears our verification rule with probability at most $2.1\times10^{-3}$. On LibriSpeech against WavMark, AudioSeal, and Timbre, KeyBound holds 1.00 detection accuracy and 0.98 bit accuracy under a spectral denoiser that costs every baseline its detection, decodes at chance without the key, and rejects transplanted carriers. Detection further transfers to held-out DAC and BigVGAN re-synthesis, though exact payload recovery degrades. Speech provenance is thus better posed as a keyed, host-bound attribution problem than as the recovery of a payload under a catalog of signal distortions fixed in advance.

View source

Similar papers

Preprint Sep 2026

DeMark: A Query-Free Black-Box Attack for Quality-Preserving Audio Watermark Removal

Audio watermarking protects digital speech by embedding imperceptible signals for ownership verification and misuse tracing. However, the security of learning-based watermarking remains insufficiently understood under realistic adversarial removal, where attackers cannot access or query the watermark encoder, decoder,...

Wei-Kang Ding, Bin-Hao Ma, Han-Qing Guo et al. · 0 citations
Preprint Aug 2026

HaloMark: A Spectral Threshold for Embedding-Vector Watermarking under C2PA

HaloMark, a watermark for embedding vectors cryptographically bound to a C2PA manifest, is presented, a watermark for embedding vectors cryptographically bound to a C2PA manifest bound rigorously for linear and non-adaptive attackers and characterise empirically for the adaptive case.

T. Sharma · 0 citations
Preprint Aug 2026

SpreadMark: Robust Image Watermarking via Spread-Spectrum Embedding

SpreadMark keeps the embedded watermark imperceptible, maintaining high perceptual quality on both COCO and DIV2K, and is the only evaluated method retaining high detection under both the regeneration and the latent-space sparsification settings the authors test.

Wei Song, Yu-Xin Cao, Zhen-Chang Xing et al. · 0 citations
Preprint Aug 2026

DHMark: Public-Key Watermarking for LLM-Generated Text via Diffie-Hellman-Guided Rejection Sampling

Large language model (LLM) watermarking provides an important mechanism for tracing the provenance of generated text. Existing statistical watermarks are often effective and robust, but most of them rely on private detection keys, which centralizes verification and complicates public auditing. Recent public or publicly...

Haocheng Fu, Yuqi Qian, Luyao Wang et al. · 0 citations
Preprint Aug 2026

Asymmetric Phase Coding Video Watermarking

Existing video watermarking systems are symmetric: the party that can verify a mark holds the extractor weights or generator secret and can therefore also embed one. Benchmarks confirm the consequence, reporting that white-box forgery defeats all evaluated methods. We present a training-free video watermark that remove...

Guang Yang, Feng-Chen Liu · 0 citations
#machine learning Preprint Sep 2026

Tokens Change, Structure Endures: Spectral Watermarking for Generated Speech

Redwing is proposed, a design principle for robust token-level watermarking that generalizes to TTS models at a speech-quality cost close to that of KGW and shows that retokenization is not merely a source of noise: its transition structure can be exploited as a design principle for robust token-level watermarking.

Kanghwi Lee, Kyeongseok Jeong, Jeongmin Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.