Skip to content

Author

Yunpeng Li

We have 2 of 17 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models

Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment...

Miao Yu, Hao-Hao Huang, Luiza S. B. Yuan et al. · 0 citations
Preprint Aug 2026

Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders

Results show that SAE-based analysis can explain defense fragmentation and guide interpretable backdoor mitigation and show system?atic encoding differences: dirty-label backdoors are dominated by isolated interaction features, whereas clean-label backdoors rely more on heterogeneous mixtures of mixed and weight-modifi...

Yikun Zeng, Chenxu Niu, Wei Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.