How Reliable are Automated Jailbreak Evaluators? A Study of Human-Machine Agreement in Cybersecurity Multi-Turn LLM Attacks
Evaluating the effectiveness of multi-turn jailbreak attacks on large language models (LLMs) increasingly relies on automated evaluators, yet their reliability and agreement with human judgments in cybersecurity conversations have not been systematically examined. This paper presents a systematic human-machine agreemen...