Text-to-image (T2I) models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these capabilities. This creates a cross-modal attack surface in which harmful semantics can remain inconspicuous in a serialized...
Zhi-Yi Mou, Yao Lu, Wang-Ze Ni et al.· 0 citations
DeShortcut-Align is proposed, a shortcut-decoupling alignment framework that reduces dependence on superficial cues that significantly improves robustness against template-stripping bypass attacks, substantially reduces over-refusal, and better preserves general-purpose reasoning capabilities, thereby mitigating the al...
Qi-Rui Liu, Yi-Chen Sun, Yan Wang et al.· 0 citations
Efficient On-Policy Self-Distilled Safety Alignment (EOPSA), which concentrates computational and gradient budgets exclusively on reliably supervised, safety-critical tokens, is proposed, consistently outperforming full-token distillation baselines in both safety compliance and reasoning retention.
Qi-Rui Liu, Yi-Chen Sun, Yan Wang et al.· 0 citations
Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer review mainly focus on the overall quality of generated reviews or the accuracy of final decisions. They therefore provide limited evidence about whether model decisions are supported by sufficient and reliable revie...
Siming Yuan, Xue-Yi Zhang, Wang-Ze Ni et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.