Skip to content

Author

Qiaosheng Zhang

We have 4 of 12 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

ReCode is introduced, a compositional budget-efficient attack that combines desensitization rewriting with two effective low-cost primitives identified by Fair-ASR, to provide a comparable evaluation basis for black-box jailbreak attacks under shared target-call budgets B.

Zhida He, Xiaoyu Wen, Han Qi et al. · 0 citations
#machine learning Preprint Aug 2026

Safin-1: Safety from Within through Memory-Native State Evolution

The routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior.

Ming Zhang, Kai-Sen Yang, Shu Yu et al. · 0 citations
Preprint Aug 2026

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

This work introduces \textsc{JailbreakSkill}, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities, which packages existing attack strategies into modular, agent-ready skills that can be directly reused and adaptively selected across tasks and target models.

Xiaoyu Wen, Jiajia Li, Zhida He et al. · 2 citations

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

This work presents SkillSafetyBench, a runnable benchmark for evaluating skill-facing safety failures, and suggests that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.

Chang Jin, Anr'an W'ang, Zeming Wei et al. · 11 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.