Skip to content

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning

Jul 2026 · arXiv.org · Vol abs/2607.19408 · 0 citations · 17 references
Computer Science

TL;DR

The implication is not that $N=2$ is universally sufficient, but that small-population failure in capable-model binary ES can be an implementation artifact rather than an intrinsic population limit.

Abstract

Using Evolutionary Strategies (ES) for fine-tuning large language models is attractive because it is memory-efficient, parallel, and compatible with black-box or discrete rewards. Yet its population-size conclusions conflict sharply: fine-tuning with cross-entropy (CE) reward succeeds with $N=1$, while binary-reward training often needs $N \approx 30$. We show this gap is largely about reward design and normalization, not population size. In the capable-model regime we study, z-score advantage normalization can cause $N=2$ to fail. Disabling normalization lets binary-reward ES with $N=2$ improve on GSM8K and TREC across capable models spanning 0.5B-7B, where the normalized variant collapses or degrades. This small-$N$ risk is set by reward granularity: binary accuracy reward induces a zero-advantage probability $q$ that depends in closed form on base accuracy, batch size, and intra-pair correctness correlation; a zero-training probe on Qwen2.5-Instruct/GSM8K matches the formula with mean absolute error 0.020 across 12 configurations and finds the availability threshold $N_{\mathrm{avail}}$ to be small in this capable-model regime. The implication is not that $N=2$ is universally sufficient, but that small-population failure in capable-model binary ES can be an implementation artifact rather than an intrinsic population limit.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Revisiting scaling laws for reward optimization

Scaling laws for optimization against reward models in AI alignment have pinned down how performance depends on optimization effort---measured by a KL-divergence budget relative to a reference policy. Beyond a certain budget, over-optimization (or reward hacking) can arise: because we optimize against a proxy reward mo...

Ali Aouad, Aymane El Gadarri, V. Farias · 0 citations
#machine learning Preprint Oct 2026

Metropolis-Hastings Dominates Importance Resampling for Policy Composition

Post-training a large language model (LLM) often requires exploring trade-offs between multiple rewards, but retraining for each trade-off is expensive. Decoding-time policy composition allows these trade-offs to be adjusted by combining reward-specific policies at inference time. This composition targets a weighted pr...

A. Kurennoy, R. Yarullin, Fergal Reid · 0 citations
#artificial intelligence Preprint Sep 2026

MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning

This work proposes a general RL-based framework for Distribution Matching allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution and proposes reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justific...

Sourabh Kulkarni, Ksheeraj Sai Vepuri, B. Demir et al. · 0 citations
#machine learning Preprint Sep 2026

Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamics, Convergence Rates, and Benefits of Off-Policyness

A unified theory for RE(S) is developed that covers the full spectrum of S, and can be interpreted as a stage-wise optimization process, where each stage takes $S$ gradient steps for minimizing the Kullback-Leibler distance to a fixed reward-weighted rollout distribution.

Zhi-Wei Wang, Yan-Xi Chen, Ya-Liang Li et al. · 0 citations
#machine learning Preprint Sep 2026

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this...

Tong Zheng, Skylar Zhai, Zhan Cheng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Diffusion Reward Models

Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many way...

Xiang-Yang Wang, Bing-Xiang He, Ze-Yuan Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.