Skip to content
Conference Open access

GMM-TDQN:Two-Stage Multi-Objective Reinforcement Learning for Large-Scale Edge Server Deployment

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · 0 citations · 23 references

TL;DR

GMM-TDQN is proposed, a two-stage multi-objective reinforcement learning framework for large-scale edge server deployment that adopts a Transformer-enhanced Deep Q-Network to learn adaptive deployment policies that balance multiple objectives.

Abstract

The deployment of edge servers plays a crucial role in supporting large-scale edge computing systems, where multiple conflicting objectives—such as latency, energy consumption, load balancing, and service reliability—must be jointly optimized in complex, dynamic environments. Existing solutions often struggle to scale effectively or to balance these objectives in a unified learning framework. In this paper, we propose GMM-TDQN, a two-stage multi-objective reinforcement learning framework for large-scale edge server deployment. The first stage employs a Gaussian Mixture Model (GMM) to capture spatial and workload heterogeneity, enabling an efficient reduction of the deployment search space. Building upon this structured initialization, the second stage formulates the deployment problem as a sequential decision-making task and adopts a Transformer-enhanced Deep Q-Network (TDQN) to learn adaptive deployment policies that balance multiple objectives. Extensive experiments on real-world datasets demonstrate that GMM-TDQN consistently outperforms state-of-the-art methods, achieving reductions of 29.18% in average latency and 17.55% in energy consumption, while improving load balancing by 27.50% and system reliability by 32.55%. These results validate the effectiveness and scalability of the proposed framework for multi-objective edge server deployment.

Read PDF

Similar papers

Preprint Jun 2026

Multi-Timescale Latent-Action DRL for Joint Optimization in Edge-Cloud Networks

A two-timescale multi-layer deep reinforcement learning framework with a latent action space (2T-MDRL-LA) to jointly optimize service placement, user association, computational delegation, task offloading, and user transmit power and achieves near-optimal performance compared to branch-and-bound solutions.

V. Son, Van-Dinh Nguyen, Ngoc Hung Nguyen et al. · 0 citations
Open access Jul 2026

STQ-Scheduler: A Secure and Throughput-Aware Deep Reinforcement Learning Framework for QoE-Driven Resource Scheduling in Distributed Video Streaming Systems

STQ-Scheduler is proposed, a secure and throughput-aware deep reinforcement learning framework that integrates high-throughput data processing, Transformer-based QoE prediction, and Proximal Policy Optimization-based resource scheduling to ensure data quality and prevent data processing from becoming a bottleneck in di...

Yi-Chun Chang, Min-Wei Jiang · 0 citations
Jul 2026

Advanced deep reinforcement learning techniques for dynamic resource allocation in 5G heterogeneous networks

The proposed modified Deep Reinforcement Learning-based intelligent TDD configuration framework for adaptive radio resource allocation in 5G HetNets effectively enhances network reliability, resource utilization, and communication efficiency in dynamic 5G HetNet environments.

G. Dalton, ·. A. Bamila, Virgin Louis et al. · 0 citations
Open access Aug 2026

Multi-Objective Optimization for Data Center HVAC Systems Based on Edge–Cloud Collaborative Deep Reinforcement Learning

This paper proposes an edge-cloud collaborative physics-informed reinforcement learning framework for production data center HVAC control that integrates a physics-informed cold-start solution using Adaptive Particle Swarm Optimization, a three-time-scale edge–cloud architecture, and a constraint-aware safe projection...

Shichao Huang, Yi-Bing Zhou, Yuan Liu · 0 citations
Open access Jun 2026

Intelligent Task Scheduling in Edge-Cloud Environments Using Double Deep Q-Network Reinforcement Learning

Experimental evaluation on a heterogeneous synthetic benchmark demonstrates that the proposed DDQN scheduler reduces SLA violations by approximately 85% relative to Round Robin and 72% relative to the greedy baseline, while achieving superior energy efficiency.

Vishakha Makode, Taresh Ayaspure · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.