How Reliable are Automated Jailbreak Evaluators? A Study of Human-Machine Agreement in Cybersecurity Multi-Turn LLM Attacks
Abstract
Evaluating the effectiveness of multi-turn jailbreak attacks on large language models (LLMs) increasingly relies on automated evaluators, yet their reliability and agreement with human judgments in cybersecurity conversations have not been systematically examined. This paper presents a systematic human-machine agreement study of four automated evaluators, namely GPT-4.1, GPT-5.2, Llama Guard 3, and a rule-based classifier, using human annotations as the reference. Experiments are conducted on 762 unique adversarial dialogs per target model, Llama 2-7B and Qwen 2-7B, covering multi-turn settings, temporal prompt variations (present vs. past tense), and cybersecurity-related topics. Rather than treating classification performance alone as evidence of evaluator reliability, we distinguish inter-rater agreement from predictive performance by using Cohen's $\kappa$ as the primary agreement measure and F1-score as a complementary metric. Results show substantial agreement among human annotators, while human-machine agreement remains weak across multi-turn conditions, with $\kappa$ values below 0.36 regardless of conversation depth. Notably, GPT-based evaluators achieve moderately high F1-scores (0.42-0.72), despite their limited agreement with human judgments, demonstrating that conventional classification metrics can mask systematic evaluator disagreement. Past-tense prompt formulations further reduce human-machine agreement, decreasing $\kappa$ by 0.11-0.48 across turn depths. No evaluator achieves consistent agreement with human judgments $(\kappa \geq 0.40)$ across aggregated cybersecurity topics. These findings identify systematic limitations in current automated jailbreak evaluation and motivate evaluation protocols that combine agreement analysis, complementary performance metrics, contextual error analysis, and targeted human review for reliable multi-turn jailbreak benchmarking.