Jul 2026· Dasinya Journal for Engineering and Informatics· Vol 2· 0 citations· 37 references
TL;DR
This paper presents a framework that removes both limits by using policy-based RL, and improves on a value-based baseline (DQN) with statistical significance and performs on par with a maximum-entropy method (SAC) under the tested settings, while keeping good sample efficiency and stability.
Abstract
Self-adaptive information systems must keep their quality requirements while their environment changes at run time. Building the adaptation logic by hand is difficult, because design-time uncertainty makes it impossible to foresee every environmental change. Online Reinforcement Learning (RL) can build this logic automatically. However, the value-based RL methods used so far have two practical limits: the exploration rate must be tuned by hand, and continuous states must be discretised by hand. This paper presents a framework that removes both limits by using policy-based RL. The Analyze and Plan phases of the MAPE-K loop are redefined as a single policy-based decision step, and Proximal Policy Optimization (PPO) is applied for online adaptation in continuous and discrete action spaces. The framework is evaluated on two systems: a self-adaptive web application and a predictive process-monitoring system. Across four workload patterns and two concept drifts, the framework learns effective policies without exploration tuning or state discretisation. It improves on a value-based baseline (DQN) with statistical significance and performs on par with a maximum-entropy method (SAC) under the tested settings, while keeping good sample efficiency and stability.
Ensuring the dependable operation of modern software systems under dynamic and non-stationary operating conditions remains a major challenge in software reliability engineering. Although recent deep reinforcement learning (DRL)-based approaches have demonstrated promising capabilities for closed-loop adaptation of soft...
Sudhakar Kambhampati, G. V. Krishna· Ceylon Journal of Science· 0 citations
It is concluded that AI-enabled test automation can improve testing efficiency and adaptability when learning mechanisms are combined with controlled validation, risk-based decision criteria, and human oversight.
N. Gunasekara, Tharushi Senanayake· The American Journal of Inte...· 0 citations
This work trains software-engineering, deep-research, and general tool-use agents on gpt-oss-20b and improves task reward in all three and presents MCP-Universe RL (MCP-U RL), an open-source framework that takes over both.
Ziyang Luo, Yan Yang, Xiang-Ru Jian et al.· 0 citations
Repository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven: gains in some categories coincide with regressions in others, while aggregate resolution obscures these changes. Motivated by this category see-saw, we develop a...
Jie Zhao, Zi-Yu Jiang, Suhang Zheng et al.· 0 citations
A structured review of three key roles that RL plays in empowering OR, serving as an end-to-end solution method or as a component integrated within heuristic and exact OR methods for combinatorial optimization problems, and facilitating extended reality analysis through integration with digital twin systems is presente...
Ya-Han Lu, Dong-Yang Xia, Nurşen Aydın et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.