Skip to content

Author

Hongli Shen

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

Experiments show that with only 1K synthesized samples, AdvSafe-aligned LRMs achieve significantly stronger jailbreak robustness than existing baselines, with almost no utility degradation, demonstrating that learning unsafety knowledge enables a superior robustness-utility trade-off and generalizes beyond seen attack patterns.

Hongli Shen, Shaopeng Fu, Qinbo Zhang et al. · 0 citations