Skip to content
Conference Open access

JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models

Jul 2026 · Annual Meeting of the Association for Computational Linguistics · pp. 16006-16029 · 2 citations · ⚡ 1 influential
Computer Science

TL;DR

JailMeter is proposed, an evidence-based evaluation framework designed to more faithfully measure jailbreak effectiveness and distill into a small language model, JailMeter\textsubscript{SLM}, which maintains comparable reliability with significantly reduced computational costs.

Abstract

The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an evidence-based evaluation framework designed to more faithfully measure jailbreak effectiveness. Inspired by the Information Bottleneck theory, JailMeter applies dual-feedback optimization to filter jailbreak noise from model responses while preserving content relevant to the original malicious question. This process produces concise evidence for a rigorous assessment under which an attack is validated only when the response captures the malicious intent and delivers a complete answer, thereby signaling a substantive bypass of model safety alignment. We evaluate JailMeter on JailMeter-Eva, a challenging benchmark containing 330 human-labeled, non-rejected jailbreak instances. JailMeter achieves an accuracy of 97.27%, substantially outperforming existing evaluation methods. To support large-scale evaluation, we further distill JailMeter into a small language model, JailMeter\textsubscript{SLM}, which maintains comparable reliability with significantly reduced computational costs. Code and dataset are available at https://github.com/Magi2B0y/JailMeter.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

An Empirical Measurement of Jailbreaking Evaluators

Expert evaluation of jailbreak responses is costly and difficult to scale, so the community increasingly relies on automated evaluators to determine whether an attack succeeds. However, jailbreak studies typically validate their chosen evaluator independently, repeatedly spending resources on similar evaluation efforts...

Yujie Mu · 0 citations
Preprint Sep 2026

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

Large language models (LLMs) are vulnerable to jailbreak attacks that bypass safety alignment through carefully crafted prompts. Many existing defenses require access to model weights or internals, making them difficult to apply to black-box deployments. We propose AlcaTRAz (Anchored Tree-Rule defense Against jailbreak...

J. Res, Petr Kaska, Martin Perešíni et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Validity-Aware Jailbreak Evaluation for Large Language Models

This work proposes Sequential Epistemic and Action-Level Validation (SEAV), a verification-centric jailbreak evaluation framework that decomposes responses into ordered steps and evaluates both validity and correctness, and shows that enforcing correctness substantially reshapes measured robustness.

Qilong Wu, Sahil Wadhwa, Pranab Mohanty et al. · 0 citations
Conference Open access Sep 2026

Explaining Jailbreaks: Structured and Interpretable Safety Assessment for Large Language Models

This work proposes an explanation-aware safety framework that augments binary harmfulness detection with structured, human-interpretable explanations capturing severity, strategies, trigger spans, ratio-nales, and derived safety factors, and introduces a human–LLM hybrid annotation and canonicaliza-tion pipeline.

Sunghee Dong, Sungwon Yi, K. Bae et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.