This work presents SkillSentry, a dynamic safety-testing framework based on adaptive honey worlds, which infers the intended capability boundary of a skill, constructs an LLM-simulated environment with controlled decoy resources, and adaptively generates tasks to explore its behavioral states.
Nizhang Li, Zonghao Ying, Xiang-Fan Wu et al.· 0 citations
This paper proposes MIND, a cognitive jailbreak framework that reframes adversarial prompt generation as a belief-state inference problem over latent defense mechanisms and actively models the target system's latent defense mechanisms by interpreting multi-modal feedback as high-density signals.
Dongdong Yang, Deyue Zhang, Zhao Liu et al.· arXiv.org· 0 citations
SafeFlow is proposed, a defense framework for multi-agent systems that formalizes malicious cross-agent propagation as a semantic information-flow problem and reduces attack success rates compared to undefended baselines and external defenses while retaining high benign task completion and a high paired safe--harm success rate.
Haowen Dai, Zonghao Ying, Wenfeng Li et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.