Skip to content
Preprint

Robust Coverless Linguistic Steganography via Sentence Embedding Space with Global Resynchronization

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

A robust coverless steganographic framework that operates in the sentence embedding space rather than the token space is proposed that achieves substantial improvements in robustness, while maintaining effective embedding capacity and exhibiting strong resistance to statistical analysis.

Abstract

Linguistic steganography enables covert communication through natural language. Existing methods heavily rely on token-level operations and struggle to maintain reliability under word- and sentence-level textual perturbations. Moreover, variable-length coding-based schemes are highly susceptible to bit-slippage under minor disturbances, as perturbations cause desynchronization between embedded and extracted bit sequences. To address these issues, we propose a robust coverless steganographic framework that operates in the sentence embedding space rather than the token space. Specifically, secret messages are encoded as hierarchical clustering paths in the sentence embedding space, which enhances decoding stability against word- and sentence-level textual perturbations. To tackle the bit-slippage problem, we introduce a Global Resynchronization Mechanism (GRM) that reframes variable-length bitstreams as discrete symbols anchored to semantic subspaces, decoupling local embedding failures from global message recovery. Experimental results demonstrate that under word- and sentence-level perturbations, our approach achieves substantial improvements in robustness, while maintaining effective embedding capacity and exhibiting strong resistance to statistical analysis.

View source

Similar papers

Preprint Aug 2026

Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs

Experiments across four controlled manipulation types under ideal channel conditions show that the embedded payload, and hence an approximate reconstruction of the authentic content, is always fully recovered without bit errors, and the results indicate that the choice of neural codec is the dominant factor for detecti...

Yigitcan Özer, Xin Wang, Zhe Zhang et al. · 0 citations
Preprint Sep 2026

Feedback Coding Enables Inference-Time Covert Agentic Communication

As large language models (LLMs) are increasingly used to automate digital interactions, users can leverage LLM-generated text as cover for covert communication within seemingly benign conversations. Existing LLM steganography, however, is predominantly white-box, requiring the sender and receiver to share the cover sta...

Si-Dong Guo, Sajani Vithana, Atefeh Gilani et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CARTS: Contextual Autoregressive Rank Transcoding Steganography for Full-Capacity Keyed Text Encoding

Autoregressive language models can be used to transform a payload text into a stegotext of identical token length by preserving per-position rank information across contexts - a methodology we formalize as Contextual Autoregressive Rank Transcoding Steganography (CARTS). While the Calgacus construction of Norelli et al...

Wissam Ghantous, Alexander V. Mantzaris · 0 citations
#machine learning Preprint Sep 2026

WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading

WeaveMark, a robust and scalable multi-bit LLM watermarking scheme based on coded payload spreading, improves payload capacity through multi-bit-per-token spreading (weaving), improving extraction accuracy through soft-decision error-correcting codes, and preserving text quality through unbiased multilayer reweighting.

Gang-Hyun Park, Ju-Hyeong Lee, Heeyoul Kwak et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.