Skip to content

RLDMGO: a reinforcement learning-driven hybrid framework for complex optimization problems

Aug 2026 · Computing · Vol 108 · 0 citations · 60 references

TL;DR

This study suggests that RLDMGO can serve as a viable and adaptive solver for complex optimization problems and achieves a competitive ranking among fourteen evaluated state-of-the-art competitors.

View source

Similar papers

#reinforcement learning Open access Aug 2026

A reinforcement-learning-guided memetic Narwhal Optimization Algorithm for global and engineering optimization

The Narwhal Optimization Algorithm is a recent swarm metaheuristic that, like most population-based optimisers, is prone to premature convergence, is sensitive to random initialisation, and relies on a rigid, schedule-driven exploration–exploitation balance. This paper develops and rigorously evaluates two enhanced variants that address these weaknesses. NWOA-OBL adds opposition-based initialisation and a stagnation-triggered, dynamic-opposition restart that replenishes population diversity, while NWOA-RL replaces the fixed exploration ratio with a Q-learning controller that selects the search behaviour online from the observed progress of the optimisation. Both variants are made memetic through a shared elite local search that supplies the local-refinement drive the original wave-based moves lack. The variants are compared against the baseline algorithm and seven established and recent optimisers on the CEC2017 suite at dimension thirty and the CEC2022 suite at dimensions ten and twenty, on six constrained engineering-design problems, and through parameter-sensitivity and ablation studies, all under a common evaluation budget with thirty independent runs and full nonparametric statistical analysis. Pooled over the benchmark functions, NWOA-RL attains the joint-best mean rank, statistically indistinguishable from the strongest competitor and significantly ahead of the remaining baselines, and reaches near-optimal engineering designs. The ablation identifies the elite local search as the decisive component of the design.

A. Al Tawil, S. Z. Hashim, Hanaa Fathi et al. · 0 citations
Open access Jul 2026

AutoPSO: A Meta-Framework for Automated Particle Swarm Optimization

Comprehensive experiments on numerical benchmarks and neuroevolution robotic control tasks demonstrate that AutoPSO consistently discovers novel PSO variants that significantly outperform strong baselines and confirm that AutoPSO achieves increasing performance gains with larger swarm sizes.

Xin-Meng Yu, Jia-Xin Gao, Jianguo Zhang et al. · 1 citation
Aug 2026

Quality-Diversity Reinforcement Learning using Behavior Regulated Policy Gradient

The family of Quality-Diversity algorithms, such as MAP-Elites, aims to generate a large collection of diverse and highperforming solutions using Evolutionary Computation. Despite their success in domains like evolutionary robotics, relying heavily on random mutations inspired by Genetic Algorithms (GA) makes MAP-Elites inefficient in evolving high-dimensional solutions. This limitation motivated the creation of PGA-MAP-Elites, which incorporated policy gradient (PG) from deep reinforcement learning and enabled evolution on large neural networks. DCRL-MAP-Elites, the latest successor of PGA-MAPElites, further advanced the performance leveraging an actor-critic training that is conditioned on the descriptor. However, the critic evaluation in DCRL-MAP-Elites is made in a non-Markovian manner, which could mislead the neuroevolution with incorrect gradient signals. Therefore, we introduce BRPG-MAP-Elites, a new QD-RL algorithm adopting a strictly Markovian actor-critic architecture within the MAP-Elites framework. BRPG-MAP-Elites learns to predict the impact of actions on both behavior and fitness, using this information to create alternative descriptor-conditioned mutations. Following a comprehensive evaluation on a wide range of locomotion control tasks, our method demonstrates a 43% improvement in average QD-scores over DCRL-MAP-Elites and achieves higher robustness in the generated policies.

Runjun Mao, Antoine Cully · 0 citations
Jul 2026

MLDGWO: a grey wolf optimizer with momentum, leader adjustment, and differential perturbation for global optimization problems

Applied to five engineering optimization problems, modified momentum–leader–differential grey wolf optimizer consistently achieves the lowest objective values, while analytically proving its ability to navigate heavily penalized boundaries and satisfy all constraints.

Xin Su, Yichen Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.