Collaborative Cloud-Edge Computing via LLM-Guided Constrained Reinforcement Learning
Abstract
Cloud computing (CC) and mobile edge computing (MEC) has become essential for augmenting computational capacity and minimizing latency by offloading tasks from user equipments (UEs) to servers. While deep reinforcement learning (DRL) is a powerful tool for optimizing offloading decisions and managing resource allocations, it can suffer from suboptimal policies, slow convergence, and constraint violations as the network scales. Large language models (LLMs) with strong reasoning capabilities and extensive prior knowledge offer a potential solution by enabling more efficient exploration under LLM guidance. Therefore, we propose an LLM-guided constrained proximal policy optimization (LGC-PPO) algorithm that integrates LLMgenerated policies, composed of action demonstrations and reward shaping. A probabilistic action sampling strategy is designed to control the action selection from LLM policy and PPO agent, which facilitate learning in high-dimensional decision spaces and improve early-stage decision quality. Experimental results demonstrate that the proposed algorithm consistently outperforms baseline methods, particularly in large and complex decision spaces, highlighting its strong potential for resource optimization in collaborative cloud-edge computing networks.