This work proposes Defense-as-Skill, a defense paradigm that implements the runtime guard itself as an installable, inspectable, and editable skill, and demonstrates transfer across victim models, held-out risk families, and external benchmarks, as well as retained protection against adaptive attackers.
Xiao-Fan Yang, Ziqi Miao, Dianbo Sui et al.· 0 citations
SCHEMA reveals that hallucinations concentrate at a small set of highly connected knowledge hubs, and that final-answer accuracy decouples from trajectory honesty; models often reach correct conclusions through structurally flawed reasoning.
Xinshun Feng, Ziqi Miao, Li-Jun Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.