Skip to content

Author

Cheng-Hao Yang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

Feature Norm Normalization: A Plug-and-Play Module for Boosting Model Accuracy in Machine Unlearning

Driven by stringent privacy regulations and data deletion requirements, machine unlearning has emerged as a critical field focused on selectively removing the influence of specific data from the pre-trained model. In this paper, we focus on achieving the unlearning objective while maintaining better model accuracy. Firstly, we identify and quantitatively characterize a previously overlooked cause of model accuracy drop during the unlearning process: feature norm shift. Then, to address this, we propose a simple yet efficient plug-and-play module, namely Feature Norm Normalization (FNN). Notably, our FNN can be seamlessly integrated into existing unlearning schemes to explicitly constrain the feature norm shift, and thus stabilize model accuracy. Extensive experiments also show that FNN can effectively help existing unlearning schemes achieve higher model accuracy. For instance, on CIFAR10 with ResNet18, using Salun with FNN for randomly unlearning 64 data achieves 50.89% higher accuracy than standard Salun.

Cheng-Hao Yang, Xue Yang, Xiaohu Tang · 0 citations
#artificial intelligence Preprint Sep 2026

Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR

Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL reward, fusing the two signals within a single step: either as a \emph{weighted-additive combination} or a \emph{teacher-modulated rescaling} of the RL advantage. In this paper, we show that a simple two-stage scheme, OPD-then-RL, consistently outperforms pure OPD, pure RLVR, and all such joint baselines across logic and math reasoning benchmarks. Beyond the empirical results, we further provide a systematic understanding of this through pass@$k$ behavior, learning dynamics, and parameter updates, yielding a consistent explanation: OPD expands the student's coverage of teacher-supported solutions and RL sharpens within that support, while jointly optimizing the two signals causes them to interfere. To provide a practical recipe, we find that the OPD validation score is the key signal for when to switch to RL, and that OPD is a better cold start for RL than SFT. Together, our results establish OPD-then-RL as a simple yet strong way to combine the two methods, turning two entangled signals into complementary stages.

Bo-Yang Li, Bingsen Chen, Cheng-Hao Yang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.