Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 12 (2026), Issue 2

Linguistic Steganography with Large Language Models: Embedding Secrets in Natural Text

Authors

Sonali Yadav, Hitesh Singh, Aditee Mattoo

Abstract

Linguistic steganography (LS) refers to the act of secretly steganographically encrypting information in natural language covertext. The new possibilities in large language models (LLMs) represent recent achievements that have created opportunities to create texts of high quality that may be used as a carrier of hidden messages. The paper has given a detailed outline of a full pipeline of generative linguistic steganography based on LLMs, describing how secret bitstreams in fluent language are encoded and how authorised decoders are assured to recover them. We are surveying past methods in text-based steganography, beginning with primitive techniques of lexicon or syntactic alteration to neural methods of text generation. Our subsequent suggestion is an innovative steganographic scheme based on LLM, that makes use of probabilistic encoding of a token (through a range coding) to steganize bits of payloads into the implied word selection, and leave intact the statistical and perceptual characteristics of natural speech. It offers adjustable embedding rates and includes discourse-sensitive element in order to preserve semantic integrity and style consistency and ensure that the steganographic text (stego-text) is stenographically innocent. We carry out the prescribed system based on a refined transformer model and test its applicability on several datasets. Experimental evaluation on WikiText-103 and RealNews datasets demonstrates an average embedding capacity of 3.98 bits per word, with perplexity reduced to 59.8, and statistical divergence (JS divergence) maintained below 0.03. Steganalysis detection accur The produced stego-texts resemble regular text almost perfectly in terms of their perplexity and semantical consistency, and cannot be detected by the latest steganalysis systems (with detection rates near to 50 percent, which is comparable to random guessing). We further compare infidelity with previous work, where our approach yields higher than average improvement of 20-30 on anti-detection success, and a large secret payload capacity. Its results demonstrate that it is indeed possible to use LLMs to implement covert communication at the borderline of safe steganographic text production.