Skip to content

Efficient Long-Horizon Learning for Learned Optimization

Jul 2026 · arXiv.org · Vol abs/2607.06772 · 0 citations · 43 references
Computer Science

TL;DR

Efficient Long-hOrizon (ELO) learning is introduced, an efficient meta-training algorithm that reallocates wasted meta-training compute to longer failure regimes, achieving efficient long-horizon learning, and enforces decoupled progressive expert supervision, providing stable meta-learning signals that additionally improve the generalization of LOs.

Abstract

Learned optimization aims to improve upon hand-designed optimizers (e.g., Adam and Muon) by meta-learning small neural network optimizers over a distribution of tasks. While recent work has greatly advanced the architectural design and inductive biases of learned optimizers (LOs), their meta-training remains biased toward short-unroll learning on particular tasks, resulting in redundant computation and leaving LOs often unable to compete with hand-designed optimizers. We introduce Efficient Long-hOrizon (ELO) learning, an efficient meta-training algorithm that (1) reallocates wasted meta-training compute to longer failure regimes, achieving efficient long-horizon learning, and (2) enforces decoupled progressive expert supervision, providing stable meta-learning signals that additionally improve the generalization of LOs. Our empirical study evaluates ELO for meta-training both element-wise and matrix-based LOs. Across downstream language modeling (GPT-2-124M/350M on FineWeb) and image classification (ViT-B/16, ResNet-50 on ImageNet-1K) tasks, ELO substantially improves the long-unroll performance and out-of-distribution generalization of the base LOs. In particular, ELO-Celo2 consistently outperforms well-tuned AdamW across all evaluated tasks, while remaining competitive with Muon on language modeling. \textit{Notably, all ELO baselines require less than 7 H100 GPU-hours for meta-training.}

View source

Similar papers

#natural language process... Preprint Sep 2026

Expert-Space Exploration in MoE Reinforcement Learning

Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the s...

Hong-Yi He, Zheng-Wen Lin, Xiao Liu et al. · 0 citations
Jul 2026

FunL2O: LLM-Guided Feature Function Design for Learning to Optimize

This work introduces FunL2O, the first unified framework for automating feature design through LLM-driven program evolution for L2O, and establishes LLM-driven feature evolution as a general and effective approach to automating representation design in L2O.

Bingheng Li, Junyang Cai, Yupeng Zhang et al. · 0 citations
#machine learning Preprint Aug 2026

Elite-Weighted Supervised Fine-tuning for Goal-Directed Molecular Optimization

This work introduces Elite-Weighted Supervised Fine-tuning (EW-SFT), which uses reward to guide elite selection of high-scoring molecules, and updates the model by its own pretraining loss on that set, and consistently outperforms the corresponding native optimizers.

Shiyun Wa, Yifei Wang, A. G. Green et al. · 0 citations
Preprint Aug 2026

TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

A simple modification to supervised fine-tuning, TailSFT, which filters out already fit sequences during training, thereby focusing learning on under-modeled regions, or the tail, of the data distribution, and introduces a lightweight diagnostic for identifying settings where TailSFT is most likely to help.

Sadhika Malladi, Samy Jelassi, Dylan J. Foster et al. · 1 citation
Jul 2026

ISO: An RLVR-Native Optimization Stack

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through t...

Hanqing Zhu, Wenyan Cong, Zhizhou Sha et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.