A Critical Review of LLM Agents for Automated Penetration Testing: Benchmark Realism and Evidence Grading
Penetration testing has long resisted full automation because it requires contextual reasoning, adaptive tool use, and experience-driven decision making across the Penetration Testing Execution Standard (PTES) lifecycle. Recent advances in large language models (LLMs), agentic reasoning, and multi-tool invocation have...