Skip to content
Preprint

Many Optimizers But Only One Training Path: Repeated Resampling for Adaptive Optimizer Selection

Aug 2026 · 0 citations · 20 references
Computer Science

TL;DR

Two variants of ROR are compared on MNIST, Fashion-MNIST, and two motor insurance claim-count models to support short scouting as a practical way to search over optimizers without completing every candidate run.

Abstract

An optimizer is usually chosen before training a deep neural network and then kept fixed. Treating optimizer choice as a hyperparameter could boost performance, but it requires several complete training runs and discards all but the winner. Repeated Optimizer Resampling (ROR) instead searches during one evolving run. Every $b$ epochs, each candidate optimizer scouts from the current model weights for $s$ epochs. The best scout continues for the remaining $b-s$ epochs, and that completed segment becomes the new incumbent if it improves the validation objective. This design allows the preferred optimizer to change as training progresses. We compare two variants of ROR on MNIST, Fashion-MNIST, and two motor insurance claim-count models. Nine fixed optimizers and both ROR variants are evaluated with the same ten seeds. One-epoch ROR uses 24\% to 35\% of the aggregate training needed to identify the best fixed optimizer exhaustively and remains close to that optimizer on all four tasks. These results support short scouting as a practical way to search over optimizers without completing every candidate run.

View source

Similar papers

#machine learning Preprint Sep 2026

Reinforcement learning to choose optimizers

Reinforcement Learning to Choose Optimizers is introduced, which formulates the optimization algorithm choice as a sequential decision-making problem and outperforms every portfolio optimizer at all but the smallest budgets.

Martin P van der Schelling, D. Toshniwal, M. A. Bessa · 0 citations
#machine learning Preprint Sep 2026

EvE: An Alternate Optimizer to Adam

EvE (Evolutionary Explorer), a steady-state, population-of-four differential evolution (DE) optimizer with a targeted Adam fallback: each iteration proposes one candidate via DE, running a short burst of gradient descent only if the DE step fails to improve on the incumbent.

Shashank Raj, Kalyanmoy Deb · 0 citations
Preprint Aug 2026

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

CODS is introduced, a critic-guided selector that alternates between fitting an algorithm-matched critic and acquiring high-residual transitions before freezing a reusable subset, a reusable selection procedure, not a formal coreset guarantee.

Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan et al. · 0 citations
#machine learning Preprint Sep 2026

Backpropagated Output Momentum: Relocating Optimizer History from Parameters to Task Space

Optimizer momentum is usually stored as a parameter-sized moving average of past gradients, which makes history costly and fixes each past signal in the coordinates in which it was computed. We introduce Backpropagated Output Momentum (BOM), which instead stores a compact moving average of prediction errors at the mode...

Yu-Chen Li, Zong-Qi Fan, Nguyen H. Tran et al. · 0 citations
#machine learning Preprint Sep 2026

Every Batch Is Its Own Validation Set: Leave-One-Out Gradient Matching for Online Data Selection in LLM Fine-Tuning

Online batch selection fine-tunes a language model on the most useful part of each candidate batch. Selectors that match the gradient of the candidate batch are attractive because they need no held-out data, yet they rarely beat training on the whole batch. We show why. In-sample gradient matching uses every example as...

Hong-Yu Chen, Xin-Yi Luo, Ming Zhao et al. · 0 citations
#machine learning Preprint Sep 2026

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.

Hao Liang, Ming-Rui Chen, Hengyi Feng et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.