Skip to content

Author

Kunlong Chen

We have 3 of 18 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Hyperparameter Scaling Laws Across MoE Sparsity

This work shows that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone.

Chang-Xin Tian, Kun-Long Chen, Jia Liu et al. · 0 citations

SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling

This work proposes SuperValid, a framework that synthesizes OOD, capability-aligned validation data by distilling core concepts from benchmarks within a capability domain and expanding them into diverse, knowledge-rich texts, which enables effective model selection, early stopping, and scaling decisions.

Quan Sun, Chang-Xin Tian, Kensen Shi et al. · 1 citation · ⚡1
Preprint Aug 2026

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

Results establish ELR as a common coordinate linking LR scheduling, norm control, and loss dynamics, and Controlled interventions further show that weight decay and Hyperball shape loss dynamics primarily through the ELR schedules they induce.

Zi-Han Liu, Rui-Heng Zheng, Shaobo Zhang et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.