On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-toke...
Zhe-Xu Wang, Mao-Lin Luo, Yan-Kun Hong et al.· 0 citations
Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adapta...
Tao Hu, Zhi-Nuo Zhou, Xia-Liang Tong et al.· 0 citations
AlgoEvo is introduced, a unified agentic architecture that transforms automated algorithm discovery into an interactive, knowledge-accumulating process, demonstrating strong intra-task accumulation, cross-task transfer, and the ability to reproduce or exceed the strongest existing methods through flexible skill activat...
Jun-Hao Qiu, Qing-Long Hu, Ji Cheng et al.· 1 citation
Zeroth-order (ZO) optimization offers a memory-efficient alternative for LLM fine-tuning by estimating updates only from forward evaluations of perturbed parameters, without backpropagation or activation storage. However, in billion-parameter LLMs, isotropic perturbations often waste many forward evaluations on weakly...
Yue Xie, Zhi Zheng, Yun-Peng Ba et al.· 0 citations
Large language models have advanced automated algorithm discovery by synthesizing executable code, but existing frameworks trap them in rigid search pipelines with pre-defined control flows. This limitation restricts adaptive reasoning, blocks cross-paradigm transfer, and discards valuable execution feedback. We propos...
Jun-Hao Qiu, Qing-Long Hu, Xia-Liang Tong et al.· 0 citations
This work proposes Hyper-ES, a subspace-based ES framework that avoids the weakness of ES in full-parameter search while exploiting its strength in low-dimensional optimization, and consistently outperforms GRPO-LoRA while requiring 10% fewer space-consuming gradient updates.
Yuntian Gu, Zhi Zheng, Yun-Peng Ba et al.· 0 citations
These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO, and study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM.
Yunpeng Ba, Zhi Zheng, Yue Xie et al.· 0 citations
DyCA treats instance clustering as a co-evolving component within the search process, reusing accumulated evaluation data as feature-free signals to progressively partition instances with similar algorithmic response patterns, thereby enabling finer-grained and more adaptive guidance for specialized algorithm design.
Qinglong Hu, Qingfu Zhang, Fei Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.