Skip to content
Open access

A Fine-Grained Semantic Steganography Framework with Cross-Modal Drift Regularization

Sep 2026 · Mathematics · 0 citations · 23 references

Abstract

Steganography using deep learning can preserve pixel-level image quality while still changing object, attribute, or relational information in captions generated by vision-language models (VLMs). This caption drift creates a detection channel that is not measured by global image-embedding similarity alone. This paper presents StegoGuard, a framework that embeds secret payloads while enforcing caption-level semantic consistency. Concept drift is defined as a measurable divergence in object-level, attribute-level, or relational semantics between cover and stego captions and is quantified using CLIP text-embedding cosine distance. The main technical contribution is a cross-modal semantic drift regularization term based on BLIP-2 captions generated for cover and stego images. The framework combines this objective with CLIP-based saliency-guided region selection and a lightweight Vision Transformer encoder-decoder. Saliency-map quality is evaluated against ground-truth segmentation masks, and a deterministic bit-to-patch mapping protocol is provided for reproducibility. Experiments use COCO2017, DIV2K, and BOSSBase. Within the controlled six-baseline protocol, StegoGuard achieved a PSNR of 38.9 dB, an SSIM of 0.976 at 256 bits, detector AUC values of 0.521–0.562, and caption similarity of 0.962 as measured by CLIP text cosine similarity (model-relative, not human-verified). The ablation results show that the drift term reduces the measured concept-shift rates while preserving the reported bit-recovery and image-quality levels.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.