Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. Harmful intent is a property of the prompt. Jailbreak success is an out...
Ming-Yu Luo, Ming Deng, Zilang Qiu et al.· 0 citations
Wrong-result bugs undermine database reliability by allowing queries to com-
plete successfully while silently returning incorrect data. Cross-product diversity can help
expose these failures, but requires replicated state and compatibility across database sys-
tems. MimicDB explores an alternative: selective query pro...
Ze-Lin Wang, Ping Chen, Jin Wei et al.· Security and Safety· 0 citations
ActProbe, an internal-state-based framework for detecting and localizing poisoned segments in multi-source LLM inputs, is proposed and remains effective against defense-aware adaptive attacks and can protect black-box APIs through surrogate-based poisoned-segment removal.
Xue Tan, Chang-Hui Wang, Sanrui Yang et al.· 0 citations
AcMAS is proposed, an activation-based framework for malicious-behavior detection in MAS that significantly outperforms graph-based baselines against stealthy attacks, with generalization across diverse open-source LLM backbones, attack intensity, and MAS scale.
Haowen Xu, Xue Tan, Lei Ma et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.