Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection
An exploratory case study that applies explainable artificial intelligence techniques to analyze how Prompt Guard 2 distinguishes malicious from benign prompts finds that Prompt Guard 2's decisions rely on the cumulative contribution of many tokens rather than a few dominant ones, yet saliency-guided synonym substituti...