Skip to content

Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection

Sep 2026 · 0 citations · 19 references
Computer Science

TL;DR

An exploratory case study that applies explainable artificial intelligence techniques to analyze how Prompt Guard 2 distinguishes malicious from benign prompts finds that Prompt Guard 2's decisions rely on the cumulative contribution of many tokens rather than a few dominant ones, yet saliency-guided synonym substitution and sentence-level paraphrasing can flip its predictions.

Abstract

Large language models (LLMs) are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation through prompt injection and jailbreak attacks. Classifier-based guardrails, such as Prompt Guard 2, are widely used as a first line of defense against such attacks, but their internal decision logic is largely opaque to both defenders and attackers. This paper presents an exploratory case study that applies explainable artificial intelligence (XAI) techniques to analyze how Prompt Guard 2 distinguishes malicious from benign prompts. We conduct four experiments to probe this question empirically. Guided by Vanilla Gradient and SHAP attributions, we find that Prompt Guard 2's decisions rely on the cumulative contribution of many tokens rather than a few dominant ones, yet saliency-guided synonym substitution and sentence-level paraphrasing can flip its predictions while altering only a moderate fraction of the text, in some cases yielding a successful jailbreak against the underlying LLM. A dataset-scale saliency analysis further shows that undetected injection prompts systematically lack the lexical markers the classifier relies on. We discuss the implications of these findings for the design and evaluation of classifier-based guardrails, and argue that explanation methods intended to support transparency can simultaneously lower the cost of constructing successful adversarial bypasses.

View source

Similar papers

Preprint Aug 2026

Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifolds

AEGIS (Adaptive Ensemble Guard for Injection Shielding) extracts instruction-sensitive projectors to identify malicious instructions and leverages a Unified Multi-Layer Consensus mechanism that aggregates topologically distinct signals across the network depth.

Jia-Hao Chen, Ruiping Yin, Xin-Feng Li et al. · 1 citation · ⚡1
#artificial intelligence Preprint Oct 2026

BRANCH: Bypassing Multi-Scanner AI Guardrails

AI systems increasingly rely on Large Language Models (LLMs) as core reasoning engines, making them targets for prompt injection and jailbreaks. Guardrails monitor and validate model inputs and outputs, yet their isolated, task-focused detection leaves gaps in their classification making them susceptible to bypasses. I...

William Hackett, Peter Garraghan · 0 citations
#artificial intelligence Preprint Sep 2026

SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement

Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has revealed that LLMs remain vulnerable to adversarial jailbreak attacks that can bypass safety guardrails and elicit harmful responses. Many defense methods are proposed to dete...

Quoc le Viet Vo, Trung Le, D. Ranasinghe et al. · 0 citations
Preprint Sep 2026

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

AlcaTRAz (Anchored Tree-Rule defense Against jailbreaks), a prompt-level defense based on rule trees that operates exclusively on the input text and requires no modification or retraining of the target model, is proposed.

J. Res, Petr Kaska, Martin Perešíni et al. · 0 citations
#machine learning Preprint Oct 2026

Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild

LLM agents increasingly rely on activation probes as runtime monitors for prompt injection, jailbreaks, and unsafe requests, reading the model's own hidden state to catch a harmful input before the agent acts on it. A cheap, increasingly common move, borrowed from LLM-as-judge prompting, is to append a short classifica...

Elad David, M. Fomin · 0 citations
Open access Dec 2025

AI security beyond core domains: resume screening as a case study of adversarial vulnerabilities in specialized LLM applications

A benchmark for this vulnerability in LLM-based resume screening is introduced: 463 job-candidate pairs drawn from a 14-domain corpus, with the evaluated sample covering 13 domains, attacked through a taxonomy of four attack types and four injection positions.

Hong-Lin Mu, Jinghao Liu, Kai-Yang Wan et al. · 4 citations · ⚡1

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.