Deep Reinforcement Learning Based Multi-Objective Auto-Scaling for Energy-Efficient Hybrid Cloud Infrastructures
Abstract
Cloud computing infrastructures use lot of energy, accounting for significant portions of global carbon emissions and operation expenses. This paper introduces a deep reinforcement learning (DRL) approach for multi-objective auto-scaling in a hybrid cloud system to optimise energy usage and service level agreement (SLA) compliance. We apply and benchmark three SOTA DRL algorithms: Proximal Policy Optimization (PPO), Deep Q-Network (DQN), Asynchronous Advantage Actor-Critic (A3C). The system model represents auto-scaling as a multi-objective reward function (MDP) between energy efficiency, SLA violations and resource consumption. The Alibaba Cluster Trace v2021 is a dataset of 11.7 million task events over 4,000 machines over 30 days, used for training and evaluation, with CloudSim-Plus simulation used to conduct controlled experiments. Experimental results show that the proposed PPO-based agent can reduce the energy consumption by 25% under 2.5% SLA violation-rate while DQN reduces the energy consumption by 18% under 8% SLA violation-rate and research of A3C reduces the energy consumption by 20% under 6% SLA violation-rate. PPO can save 32% in operational costs with an average CPU utilization of 82.4% compared to static threshold baselines. This framework directly contributes to the achievement of SDG 7 (Affordable and Clean Energy), SDG 9 (Industry, Innovation and Infrastructure) and SDG 13 (Climate Action) by providing data centers with a way to substantially cut carbon emissions without compromising the quality-of-service guarantees. The outcomes prove that the PPO algorithm is the best option for hybrid cloud auto-scaling, providing an energy-efficient and scalable solution for managing cloud resources.