This work shows that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone.
Chang-Xin Tian, Kun-Long Chen, Jia Liu et al.· 0 citations
This work proposes SuperValid, a framework that synthesizes OOD, capability-aligned validation data by distilling core concepts from benchmarks within a capability domain and expanding them into diverse, knowledge-rich texts, which enables effective model selection, early stopping, and scaling decisions.
Quan Sun, Chang-Xin Tian, Kensen Shi et al.· arXiv.org· 1 citation· ⚡1
Results establish ELR as a common coordinate linking LR scheduling, norm control, and loss dynamics, and Controlled interventions further show that weight decay and Hyperball shape loss dynamics primarily through the ELR schedules they induce.
Zi-Han Liu, Rui-Heng Zheng, Shaobo Zhang et al.· 3 citations
This work presents a stable and efficient training pipeline, incorporating algorithmic and system optimizations such as clipped importance sampling, training-inference ratio correction, and mixed-precision control, and proposes a structured evaluation framework across three dimensions: comprehensibility, reproducibilit...
Xinyu Tang, Gangqiang Cao, Yurou Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.