Skip to content

MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment

Jul 2026 · arXiv.org · Vol abs/2607.11070 · 0 citations · 64 references
Computer Science

Abstract

Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, yet existing learning signals are often too coarse to identify their individual contributions. We propose decomposed credit GRPO (DC-GRPO), a unified turn-level credit assignment framework for Group Relative Policy Optimization in multi-turn jailbreak learning. DC-GRPO assigns a separate group-relative learning signal to each turn by combining immediate and future credit, avoiding the credit misassignment induced by broadcasting a single trajectory-level score across the dialogue. We instantiate this framework with static and dynamic weighting rules that differ in how the two credit sources are balanced while sharing the same turn-level structure. Across multiple victim LLMs and benchmarks, the dynamic- and static-weighted variants achieve average ASR5@3 scores of 98.26% and 97.88%, respectively, substantially outperforming the state-of-the-art methods, including SEMA (86.58%) and TROJail (86.23%). Their consistently strong performance indicates that the central empirical benefit comes from turn-level group-relative credit assignment rather than a particular weighting rule. Warning: This paper contains examples of harmful content.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection

Overall, how agents are differentiated produces larger performance shifts than topology, which should be evaluated jointly with specialization: how agents are differentiated produces larger performance shifts than topology, which should be evaluated jointly with specialization.

Ulugbek Shernazarov, Charitha Ruwansiri Weerakon Basnayake, Abdelkhaleq El Jarjini et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations

This paper proposes SWRouter, a Similarity-Contractive Window Router for multi-turn large language model routing that combines a similarity-based context segmentation mechanism for prompt construction with a dual-metric evaluation framework that decouples construction accuracy from router performance.

Yu Wang, Yuchen Li, Rui Kong et al. · 0 citations
Preprint Aug 2026

TCPO: Turn-Level Credit Policy Optimization

TCPO casts credit assignment as score-to-credit conversion and constructs turn-level advantages through reference-based comparisons and achieves the best or tied-best best-turn Pass@8 on Qwen3-4B and DeepSeek-R1-Distill-Llama-8B, reduces turns to success, and improves multi-turn agent performance.

Si-Cong Liao, Zhi Chen, Yao-Hua Tang · 3 citations
Preprint Aug 2026

IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents

Influence-Aware Policy Optimization (IAPO), which represents each rollout as a typed influence-dependency graph over trainable agent actions, with user and tool observations serving as evidence, is introduced and advances the understanding of credit assignment in multi-turn user interactions.

B. Ren, Yirong Mao, Yi Yang et al. · 1 citation
Preprint Aug 2026

Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents

Experiments on WebShop, ALFWorld, and visual Sokoban show consistent improvements across language and vision-language models, while diagnostic ablations support the effectiveness of Bellman fixed-point value estimation and show that step-level credit should be incorporated selectively rather than uniformly into the fin...

Hongxi Yan, Ziyue Huang, Shichao Fan et al. · 1 citation
Book Open access Aug 2026

CES: Combinatorial Experts Selection via Contextual Linear Bandits

With the rapid advancement of large language models (LLMs), multi-agent systems have emerged as a promising alternative to scaling up a single model. Existing approaches ensemble multiple LLMs to improve response quality, but they often rely on static prior knowledge of model capabilities and prompts, and require exten...

Jinkun Xu, Minghan Wang, Zhiyong Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.