A robust coverless steganographic framework that operates in the sentence embedding space rather than the token space is proposed that achieves substantial improvements in robustness, while maintaining effective embedding capacity and exhibiting strong resistance to statistical analysis.
Abstract
Linguistic steganography enables covert communication through natural language. Existing methods heavily rely on token-level operations and struggle to maintain reliability under word- and sentence-level textual perturbations. Moreover, variable-length coding-based schemes are highly susceptible to bit-slippage under minor disturbances, as perturbations cause desynchronization between embedded and extracted bit sequences. To address these issues, we propose a robust coverless steganographic framework that operates in the sentence embedding space rather than the token space. Specifically, secret messages are encoded as hierarchical clustering paths in the sentence embedding space, which enhances decoding stability against word- and sentence-level textual perturbations. To tackle the bit-slippage problem, we introduce a Global Resynchronization Mechanism (GRM) that reframes variable-length bitstreams as discrete symbols anchored to semantic subspaces, decoupling local embedding failures from global message recovery. Experimental results demonstrate that under word- and sentence-level perturbations, our approach achieves substantial improvements in robustness, while maintaining effective embedding capacity and exhibiting strong resistance to statistical analysis.
(k)-SwordStamp is designed: semantic watermarks with order-robust detection over sub-sentence units, reducing sensitivity to attacker-chosen structure at a small quality cost.
Abdulrahman Diaa, Jonathan Petit, Florian Kerschbaum· 0 citations
Experiments across four controlled manipulation types under ideal channel conditions show that the embedded payload, and hence an approximate reconstruction of the authentic content, is always fully recovered without bit errors, and the results indicate that the choice of neural codec is the dominant factor for detecti...
Yigitcan Özer, Xin Wang, Zhe Zhang et al.· 0 citations
As large language models (LLMs) are increasingly used to automate digital interactions, users can leverage LLM-generated text as cover for covert communication within seemingly benign conversations. Existing LLM steganography, however, is predominantly white-box, requiring the sender and receiver to share the cover sta...
Si-Dong Guo, Sajani Vithana, Atefeh Gilani et al.· 0 citations
Autoregressive language models can be used to transform a payload text into a stegotext of identical token length by preserving per-position rank information across contexts - a methodology we formalize as Contextual Autoregressive Rank Transcoding Steganography (CARTS). While the Calgacus construction of Norelli et al...
Wissam Ghantous, Alexander V. Mantzaris· 0 citations
The survey aims to serve as both a reference and a roadmap for practical and responsible linguistic steganography in the LLM era by identifying five specific paradigm shifts in the LLM era.
Ruiyi Yan, Chenhui Chu, Zhong-Liang Yang et al.· 4 citations
WeaveMark, a robust and scalable multi-bit LLM watermarking scheme based on coded payload spreading, improves payload capacity through multi-bit-per-token spreading (weaving), improving extraction accuracy through soft-decision error-correcting codes, and preserving text quality through unbiased multilayer reweighting.
Gang-Hyun Park, Ju-Hyeong Lee, Heeyoul Kwak et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.