Jailbreak Context Lingers: Divergent Safety Routing and Its Cross-Task Predictability in Tool Agents
It is found that identical safety feedback induces sharply model-dependent behavioral routing rather than uniform protection, establishing a mechanistic lens and a predictive baseline for anticipating the safety and utility trade-offs of post-jailbreak feedback in autonomous agents.
Xi Wang, Song-Lei Jian, Yi-Ming Zhang et al.
· 0 citations