Skip to content
Open access

Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM Agents

Oct 2026 · Vol 3, pp. 221 - 245 · 0 citations · 89 references

Abstract

LLM agents that invoke external tools face critical safety vulnerabilities when malicious manipulations exploit their implicit trust in tool outputs and metadata. However, identifying these vulnerabilities through testing is challenging due to the need to bypass safety guardrails with semantically legitimate inputs, the difficulty in characterizing successful attacks amid implicit tool trust, and the requirement to maintain logical consistency across fragile state-dependent execution chains. In this paper, we first conduct an empirical study to investigate how external tools influence agent reasoning. Guided by the findings, we propose Datura, an automated red teaming testing framework that exposes safety vulnerabilities through chained tool manipulation. Through a five-stage workflow, Datura dynamically generates test cases where each individual step appears legitimate yet collectively leads to harmful outcomes. We evaluate Datura across five LLMs and 740 safety-critical tasks under five defense settings, including real-world safety mechanisms. Under Model Alignment, Datura achieves 94.86--99.59% attack success rate (ASR), outperforming the strongest baseline by up to 25.27 percentage points. Under Prompt Refuge, Datura maintains 78.78--95.54% ASR, showing that progressive tool-chain manipulation remains effective even under prompt-level safeguards.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.