Skip to content
Open access

Evil-AI Benchmark: An Evaluation Framework for Adversarial Security Assessment of LLM Agents in Smart Environments

2026 · IEEE Access · Vol 14, pp. 133098-133116 · 0 citations · 71 references

Abstract

Smart environments increasingly integrate artificial intelligence (AI) to process sensor data, coordinate Internet of Things (IoT) devices, and actuate cyber-physical services. Within these environments, large language models (LLMs) are emerging as decision-making and interaction components. However, LLM integration expands the attack surface of these environments, and existing benchmarks focus largely on risks associated with text-based interactions and LLM-generated outputs, failing to capture cyber-physical risks that arise when models control physical actuators. This paper introduces Evil-AI Benchmark, an open-source evaluation framework for assessing LLM agents in smart environments through (i) five adversarial threat vectors—prompt injection, persuasion, attacker-in-the-middle (AITM) simulation, data leakage, and unsafe actions—and (ii) a capability validation that checks whether a model can correctly activate device-control tools. The benchmark instruments tool activations and forwards tool-trigger events to an external Arduino-driven servo actuator, providing physical verification of actuation events. To demonstrate the benchmark in practice, we evaluate eight representative LLMs using 350 executable scenarios: 250 adversarial tests spanning five smart-environment domains (smart home, healthcare IoT, industrial control, public infrastructure, and smart building), 50 capability checks, and 50 benign over-refusal tests. Responses are scored by two independent judge models and verified by a human evaluator. We report Evilness as the extent to which attacks succeed. The Evilness Score equals the number of successful attacks, and the Evilness Rate is the corresponding percentage. Across models, Evilness Rates range from 0.8% to 54.0% (2–135 successful attacks out of 250), while Defense Rates range from 46.0% to 99.2%. The most persistent weaknesses occur in persuasion, especially under sustained multi-turn pressure, and in prompt injection. Moreover, larger models are not necessarily more robust against adversarial attacks. We additionally quantify over-refusal on benign authorized requests, so that safety is not rewarded by indiscriminate refusal. Evil-AI Benchmark provides reproducible diagnostics and quantitative metrics that connect text-based safety evaluation with real-world cyber-physical operation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.