Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models

Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment...

Miao Yu, Hao-Hao Huang, Luiza S. B. Yuan et al. · 0 citations
#machine learning Preprint Sep 2026

RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning

Large language models (LLMs) exhibit strong general capabilities that mechanistic interpretability has attributed to sparse computational circuits. However, existing circuit studies emphasize preserving functionality or explaining safety, leaving the mechanisms underlying failures across a broader range of tasks largel...

Chuan-Pu-Zou-Sheng-Li-Guo-Jing-Zhong-Gu-Yue-Shu-Ca Liu, Miao Yu, Yi-Kai Cai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.