Skip to content

Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPO

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

A transformer-enhanced proximal policy optimization (PPO) framework that enables efficient collaboration among MEC servers and significantly outperforms conventional PPO and heuristic-based approaches in terms of task completion rate and overall system efficiency.

Abstract

This paper investigates collaborative mobile edge computing (MEC) servers for large language model (LLM) inference under soft deadline constraints. In this system, to improve the quality of service, computations are expected to be completed within their deadlines. However, due to dependencies among tasks or subtasks, any missed deadline can lead to catastrophic consequences for the entire request. In this context, this work proposes an extended deadline mechanism with constrained flexibility. The main challenges lie in handling large-scale computations under strict latency constraints while limiting the number of allowable deadline extensions, especially in the presence of task dependencies within each request. To tackle these challenges, we develop a transformer-enhanced proximal policy optimization (PPO) framework that enables efficient collaboration among MEC servers. The proposed approach aims to maximize the number of tasks completed within their deadlines while minimizing the use of deadline extensions. By capturing temporal dependencies and cross-server interactions, the transformer improves decision-making for task migration. Simulation results demonstrate that the proposed method significantly outperforms conventional PPO and heuristic-based approaches in terms of task completion rate and overall system efficiency.

View source

Similar papers

#edge computing Open access Oct 2026

Distributional Reinforcement Learning for task offloading, resource allocation and early exit selection at the edge

Simulation results confirm that integrating EE with edge computing significantly improves the trade-off between inference accuracy and latency, achieving up to 212% improvement in the average task completion ratio compared to edge computing systems without EE, under the considered simulation settings.

Simone Angelucci, R. Valentini, M. Levorato et al. · 0 citations
#machine learning Preprint Aug 2026

DART-FL: Burst-Aware Multitask Federated Learning under Dynamic Inference Demand at the Edge

Results show that DART-FL dynamically adapts the inference-training resource split to time-varying inference demand and shifts the learning progress of high-demand tasks toward their burst periods, improving model accuracy when those tasks are frequently requested while maintaining comparable long-term multitask perfor...

Yi-Ming Xie, Pinrui Yu, Geng Yuan et al. · 0 citations
#large language models Book Open access Sep 2026

Online Scheduling of Battery-Aware Speculative Decoding for Energy-Efficient Cloud-Edge Collaborative LLM Inference

While distributed speculative decoding can offer efficient acceleration for Large Language Model (LLM) inference in cloud-edge environments, unleashing its full potential confronts significant challenges, including complex token draft-length management, uncertain prompt arrivals and system conditions, and joint edge ba...

Heng-Di Wang, Lei Jiao, Kong-Lin Zhu et al. · 0 citations
Open access Sep 2026

HRL-TaskOpt: A Hierarchical Reinforcement Learning-Based Task Scheduling Framework for Multi-Cloud and Hybrid Environments

Cloud computing has emerged as a new paradigm, which entrusts task scheduling to ensure the satisfaction of stringent constraints on latency, energy, and resources for sustainably running real-time applications. State-of-the-art natural DRL-based scheduling solutions mainly rely heavily on DRL techniques and are either...

Krishna Patwari, Raghvendra Kumar, J. Sastry · 0 citations
#edge computing Preprint Sep 2026

Iapetus: Content-Aware Hierarchical Scheduling for Collaborative ViT Inference in LEO Satellite Networks

sys is presented, a content-aware hierarchical scheduler that screens constellation-wide options to retain a bounded candidate set, then refines each candidate into a complete token compression and layer offloading trajectory using quality prediction and joint planning and balances per-task latency, energy, and quality...

Yan Chen, Yun-Xiang Zhang, Guan-Jun Jiang et al. · 0 citations
Sep 2026

Hierarchical risk-aware RL for DVFS-enabled IoE scheduling

The ER-RL-DVFS framework combines an energy-risk task-priority score, DVFS-aware cluster–PM–VM profiling, hierarchical candidate filtering, deterministic score-based assignment for clearly separated candidates, selective tabular reinforcement learning as the prescribed refinement mechanism for near-tie decisions, and m...

Amir Javadpour, Forough Ja'fari, T. Taleb · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.