Skip to content
Preprint

Robust Coverless Linguistic Steganography via Sentence Embedding Space with Global Resynchronization

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

A robust coverless steganographic framework that operates in the sentence embedding space rather than the token space is proposed that achieves substantial improvements in robustness, while maintaining effective embedding capacity and exhibiting strong resistance to statistical analysis.

Abstract

Linguistic steganography enables covert communication through natural language. Existing methods heavily rely on token-level operations and struggle to maintain reliability under word- and sentence-level textual perturbations. Moreover, variable-length coding-based schemes are highly susceptible to bit-slippage under minor disturbances, as perturbations cause desynchronization between embedded and extracted bit sequences. To address these issues, we propose a robust coverless steganographic framework that operates in the sentence embedding space rather than the token space. Specifically, secret messages are encoded as hierarchical clustering paths in the sentence embedding space, which enhances decoding stability against word- and sentence-level textual perturbations. To tackle the bit-slippage problem, we introduce a Global Resynchronization Mechanism (GRM) that reframes variable-length bitstreams as discrete symbols anchored to semantic subspaces, decoupling local embedding failures from global message recovery. Experimental results demonstrate that under word- and sentence-level perturbations, our approach achieves substantial improvements in robustness, while maintaining effective embedding capacity and exhibiting strong resistance to statistical analysis.

View source

Similar papers

Preprint Aug 2026

Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs

Partial manipulation of speech recordings, where only localized segments of an utterance are altered, poses a significant challenge for content integrity verification, as reliable detection and localization of such edits becomes harder as the manipulated proportion decreases. Watermarking offers a proactive defense alt...

Yigitcan Özer, Xin Wang, Zhe Zhang et al. · 0 citations
Preprint Sep 2026

Feedback Coding Enables Inference-Time Covert Agentic Communication

As large language models (LLMs) are increasingly used to automate digital interactions, users can leverage LLM-generated text as cover for covert communication within seemingly benign conversations. Existing LLM steganography, however, is predominantly white-box, requiring the sender and receiver to share the cover sta...

Si-Dong Guo, Sajani Vithana, Atefeh Gilani et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CARTS: Contextual Autoregressive Rank Transcoding Steganography for Full-Capacity Keyed Text Encoding

Autoregressive language models can be used to transform a payload text into a stegotext of identical token length by preserving per-position rank information across contexts - a methodology we formalize as Contextual Autoregressive Rank Transcoding Steganography (CARTS). While the Calgacus construction of Norelli et al...

Wissam Ghantous, Alexander V. Mantzaris · 0 citations
#machine learning Preprint Sep 2026

WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading

Multi-bit watermarking for large language models enables content source tracing by embedding user-identifiable messages into generated text. Existing methods face a fundamental trade-off among extraction accuracy, text quality, and payload capacity. We propose WeaveMark, a robust and scalable multi-bit LLM watermarking...

Gang-Hyun Park, Ju-Hyeong Lee, Heeyoul Kwak et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.