Skip to content
Conference

Semantic Consistency Drift Monitoring for Training-Free Adversarial Prompt Injection Detection in Agentic LLM Pipelines

Aug 2026 · International Conference Computational Vision and Bio Inspired Computing · pp. 1444-1449 · 0 citations · 21 references

Abstract

Large language model (LLM) agents that autonomously retrieve external content and execute multi-step action plans are increasingly deployed in enterprise and safetycritical settings. This architecture exposes a critical attack surface: adversarial content embedded within retrieved documents can silently redirect agent behavior away from the original user intent—a threat known as prompt injection. Existing defenses rely on keyword filtering, fine-tuned classifiers, or full model access, making them brittle against semantically obfuscated payloads and costly to deploy. We propose the Semantic Consistency Drift Monitor (SCDM), a training-free, model-agnostic detection framework that measures the cosine distance between latent semantic representations of the agent's pre-retrieval goal statement and its post-retrieval action plan. A statistically significant divergence beyond a calibrated threshold signals a potential injection. We evaluate SCDM on a controlled corpus of 40 samples spanning 15 operational domains across 20 adversarial scenarios. SCDM achieves an ROC-AUC of 0.7950, an F1-score of 0.7907, precision of 0.7391, and recall of 0.8500, with an inference latency of 1.22 ms per decision. Statistical analysis confirms strong class separation (Cohen's $d=1.19, p<0.001$). SCDM outperforms Jaccard bag-of-words (AUC $=0.715$), raw TF-IDF cosine $(\mathbf{A U C}=\mathbf{0. 7 7 3})$, and random baselines (AUC $=\mathbf{0. 5 4 5}$), while requiring no model retraining, no API access, and no labeled injection corpus. These results establish semantic consistency probing as a practical, lightweight safety primitive for agentic AI deployments.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.