2026· IEEE Transactions on Communications· Vol 74, pp. 11519-11533· 0 citations· 50 references
Computer Science
Abstract
Semantic-aware edge computing has exhibited tremendous potential for reducing communication-computing-caching (3C) resource costs in vehicular networks through task-oriented semantic extraction. However, environmental dynamics and uncertainties across heterogeneous timescales pose critical challenges for 3C resource allocation in semantic-aware vehicular edge computing (VEC) networks. To this end, this paper investigates a joint semantic content caching and semantic task offloading problem in twin-timescale semantic-aware VEC scenarios, aiming to maximize the long-term utility tradeoff between task execution latency and semantic cache hit ratio. First, twin-timescale Decentralized Partially Observable Markov Decision Processes (Dec-POMDPs) are established, where semantic content caching is optimized on a large timescale, while semantic offloading and bandwidth allocation policies are learned on a small timescale. Subsequently, a novel twin-timescale 3C resource allocation solution based on multi-agent graph reinforcement learning method is proposed. Specifically, a Graphical Partial Reward Decoupling-aided Multi-Agent Proximal Policy Optimization (GPRD-MAPPO) algorithm is proposed, which incorporates graph attention networks (GAT) and credit assignment mechanism to decouple the irrelevant agents in cooperative learning by dynamically identifying inter-agent graphical dependencies. Our simulation results verify the superiority of the proposed solution in reducing task execution latency and improving semantic cache hit ratio over the benchmarks in varying numbers of vehicles and diverse task characteristics.
A constraint-aware multi-agent edge collaborative offloading algorithm (CARE-CTDE) that achieves better scheduling performance, resource utilization, and constraint satisfaction than baseline methods in dynamic heterogeneous MEC scenarios, demonstrating its effectiveness and robustness for constrained edge computing systems.
Yuxuan Yang, Hexing Wang, Yang Zhou· Mathematics· 0 citations
6G vehicular services, including cooperative perception, augmented reality navigation, and high-definition map updating, need computation support close to moving vehicles. Vehicular Edge Computing (VEC) is a natural solution, but the offloading decision becomes difficult when wireless channel conditions, vehicle density, and edge server loads vary simultaneously. In this paper, we study joint task offloading and resource allocation in 6G VEC with high- and low-frequency cooperation (HL-FC). We formulate the problem as a decentralized partially observable Markov decision process (Dec-POMDP). Each vehicle decides its offloading ratio, transmission power, server association, and edge CPU request from local observations. To evaluate the proposed policy, we build a lightweight equation-driven Python simulator and compare MAPPO with Local-only, Edge-only, Random, and Greedy policies. Compared with Edge-only, MAPPO reduces the average system cost by 32.15%, 23.51%, and 17.13% under 10, 15, and 20 vehicles, respectively. It also improves the task completion rate by 21.00, 20.49, and 17.65 percentage points. Additional blockage experiments show that HL-FC keeps the policy more robust than high-frequency-only transmission under severe high-frequency blockage. The results reveal that MAPPO delivers better performance when edge resources become congested than in lightly loaded scenarios.
Zi-Heng Gu· 2026 8th International Confe...· 0 citations
This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.
Youssef Oukissou, Mohamed Amine Meddaoui, Ayoub Belaidi et al.· International journal of Com...· 0 citations
Unmanned Aerial Vehicles (UAVs) are increasingly deployed as embodied aerial agents in low-altitude economies, forming mobile aerial edge networks that enable flexible computation offloading for vehicles. However, their limited endurance and frequent join/leave behaviours result in highly dynamic topologies, undermining long-term resource availability. Moreover, existing vehicle-centric task scheduling strategies cause resource contention and decision complexity in dense environments. To address these challenges, this paper proposes a hierarchical and scalable reinforcement learning-based scheduling framework (SkySched). In SkySched, UAVs collaboratively make deployment and task scheduling decisions. The framework consists of two tightly coupled modules. First, an adaptive UAV deployment module introduces a capability encoding mechanism that compresses heterogeneous UAV attributes into a unified one-dimensional capability index. This compact representation enables a Scalable Proximal Policy Optimization (SPPO) algorithm to efficiently coordinate UAV positioning, maximizing task coverage and sustaining network-wide computing availability under dynamic topology variations. Second, a hierarchical task scheduling module is designed, where K-means-based Roadside Unit (RSU) clustering enables vertical task offloading, while a SPPO-driven horizontal UAV-to-UAV task redistribution mechanism achieves fine-grained load balancing across the UAV swarm. Simulations demonstrate that SkySched consistently outperforms state-of-the-art methods in terms of task coverage and load fairness, validating its effectiveness as an agentic AI-driven embodied networking solution for UAV-assisted vehicular edge computing.
Meng Yi, V. Lee, Miao Du et al.· IEEE Transactions on Cogniti...· 0 citations
Vehicular edge computing (VEC) has emerged as a key paradigm to support computation-intensive and delay-sensitive vehicular applications by offloading tasks from vehicles to nearby multi-access edge computing (MEC) servers. However, in realistic urban environments, task processing performance is heavily affected by heterogeneous vehicle-MEC interactions, spatiotemporal traffic dynamics, and continuously varying vehicle populations. To address these challenges, this paper considers a traffic-aware embodied edge intelligence-enabled vehicular network (EEIVN), where edge intelligence is grounded in the physical traffic environment by integrating VLM-based semantic perception with edge decision making. Based on this architecture, we formulate a reliability-constrained delay minimization problem (RDMP) by jointly optimizing task offloading ratio, computing resource allocation, and vehicle association, while constraining the queue reliability to mitigate queue-induced tail delay. To solve the NP-hard RDMP, we propose a VLM-multi-agent proximal policy optimization (VLM-MAPPO) approach that integrates a VLM-based traffic awareness method, a vehicle-adaptive MAPPO algorithm, and a vehicle association scoring and selection mechanism. Extensive simulations based on SUMO and CARLA demonstrate that the proposed VLM-MAPPO approach outperforms benchmarks in terms of task completion delay and tail delay, while maintaining comparable vehicle energy consumption and exhibiting robust scalability under dynamic traffic conditions and varying vehicle densities.
Xulong Qiao, Jian Wang, Zemin Sun et al.· IEEE Transactions on Cogniti...· 0 citations
With the rapid growth of Vehicular Edge Computing (VEC) and Mobile Edge Computing, efficient task offloading is essential for enhancing the computing and communication capabilities in vehicular networks. However, many existing methods suffer from slow convergence, load imbalance, and instability in dynamic, latency-sensitive environments. To address these challenges, we propose MAPPO-Lyapunov (MAPPO-L), a multi-agent offloading framework that integrates Multi-Agent Proximal Policy Optimization (MAPPO) with Lyapunov optimization. MAPPO-L enables distributed coordination among vehicles, roadside units (RSUs), and cloud servers, minimizing delay, improving resource utilization, and ensuring long-term stability. Lyapunov theory transforms long-term stability into per-slot optimizations, while MAPPO ensures efficient policy learning. An adaptive exploration mechanism dynamically adjusts exploration rates based on network dynamics, accelerating convergence and stabilizing training. Extensive simulations with real-world data show that MAPPO-L maintains task completion rates above 80%, converges 25%–37.5% faster than baselines, and reduces training fluctuations to 2.3%. Ablation studies confirm the critical roles of location, channel, and queue information, validating the robustness of MAPPO-L in practical VEC environments.
Lu Wei, Yong Yu, Jie Cui et al.· IEEE Transactions on Network...· 0 citations