Skip to content
Preprint

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging

Aug 2026 · 0 citations
Computer Science

TL;DR

This work proposes Hyper-ES, a subspace-based ES framework that avoids the weakness of ES in full-parameter search while exploiting its strength in low-dimensional optimization, and consistently outperforms GRPO-LoRA while requiring 10% fewer space-consuming gradient updates.

Abstract

Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning. However, directly applying ES to billion-parameter LLMs is highly ineffective. In such high-dimensional parameter spaces, most random perturbations are nearly orthogonal to useful update directions, leading to unstable optimization. We propose Hyper-ES, a subspace-based ES framework that avoids the weakness of ES in full-parameter search while exploiting its strength in low-dimensional optimization. Instead of asking ES to discover useful directions from random perturbations in the LLM parameter space, Hyper-ES first performs a small number of inexpensive gradient-based fine-tuning runs to obtain descent directions. Although each direction may provide only a limited improvement on its own, their span forms a compact adaptation subspace that captures useful reasoning updates. Hyper-ES then applies CMA-ES to optimize layer-wise DARE-TIES merging coefficients within this subspace, allowing ES to search over combinations of meaningful descent directions rather than over arbitrary full-model perturbations. We evaluate Hyper-ES on three Qwen2.5-Instruct and DeepSeek-R1-Distill backbones across six mathematical reasoning datasets. Results show that Hyper-ES consistently outperforms GRPO-LoRA by 1% while requiring 10% fewer space-consuming gradient updates. Code at https://github.com/kuangrepi/Hyper-ES.

View source

Similar papers

Preprint Aug 2026

Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies

Evolution Strategies (ES), a population-based, gradient-free post-training method that optimizes directly in weight space through random perturbations, achieves consistently higher pass@k than RL and produces a broader output distribution with greater solution coverage.

Conor F. Hayes, Elliot Meyerson, Kajetan Schweighofer et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR

DATPO (Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization) integrates difficulty-adaptive tree search with a sibling-diversity advantage term, explicitly promoting semantic diversity to expand reasoning coverage during training.

Young Kyu Yu, Sanghwan Jang, Hwanjo Yu · 1 citation
#machine learning Preprint Aug 2026

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO, and study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM.

Yunpeng Ba, Zhi Zheng, Yue Xie et al. · 0 citations
#natural language process... Preprint Sep 2026

Expert-Space Exploration in MoE Reinforcement Learning

Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the s...

Hong-Yi He, Zheng-Wen Lin, Xiao Liu et al. · 0 citations
Book Open access Aug 2026

VCAgent: A Mutation-Guided Self-Reflective Agent Framework for Virtual Cell Modeling

Large Language Models (LLMs) are increasingly used for agent-based virtual cell modeling, yet existing frameworks rely on unstructured retrieval or static prompt engineering, injecting noisy evidence and wasting inference budget on redundant tool-use trajectories. We propose VCAgent, a self-evolving framework that opti...

Zhiyun Li, Rong Han, Xiao-Yong Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.