Skip to content

MERA: A Green Edge Resource Control System With Privacy-Preservation via Mean-Field Reinforcement Learning

Sep 2026 · IEEE Transactions on Knowledge and Data Engineering · Vol 38, pp. 5963-5977 · 0 citations · 47 references

Abstract

The global rollout of 5G networks has spurred the rapid deployments of edge servers for hosting latency-sensitive web applications, which improves quality of experience (QoE). However, current efforts fall short in the substantial energy costs associated with the 24/7 operation of edge servers and overlook user privacy by requiring accurate user information for service provision, eroding the sustainability of multi-access edge computing (MEC). To enhance the QoE and service performance while ensuring privacy in MEC, we systematically formulate the interaction among edge servers as a privacy-preserving experience-aware edge resource control (PEERC) problem. To address this, we conduct a global resource control and propose a collaborative resource allocation system named MERA. MERA leverages <inline-formula><tex-math notation="LaTeX">$k$</tex-math><alternatives><mml:math><mml:mi>k</mml:mi></mml:math><inline-graphic xlink:href="xia-ieq1-3705464.gif"/></alternatives></inline-formula>-anonymity data obfuscation to protect user location and resource demand privacy while enhancing service performance and energy efficiency with mean-field multi-agent reinforcement learning. Extensive experiments based on a synthetic real-world dataset demonstrate that MERA significantly surpasses benchmarks in terms of QoE, user coverage, privacy, and energy efficiency by <inline-formula><tex-math notation="LaTeX">$1.18\times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>1</mml:mn><mml:mo>.</mml:mo><mml:mn>18</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="xia-ieq2-3705464.gif"/></alternatives></inline-formula>, <inline-formula><tex-math notation="LaTeX">$1.24\times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>1</mml:mn><mml:mo>.</mml:mo><mml:mn>24</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="xia-ieq3-3705464.gif"/></alternatives></inline-formula>, <inline-formula><tex-math notation="LaTeX">$1.63\times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>1</mml:mn><mml:mo>.</mml:mo><mml:mn>63</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="xia-ieq4-3705464.gif"/></alternatives></inline-formula>, and <inline-formula><tex-math notation="LaTeX">$1.27\times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>1</mml:mn><mml:mo>.</mml:mo><mml:mn>27</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="xia-ieq5-3705464.gif"/></alternatives></inline-formula> on average.

View source

Similar papers

Open access Jul 2026

Hierarchical Mean-Field Theory-based Off-Policy GRPO for Federated Edge Learning in Resource-Constrained Edge Computing

With the popularization of Internet of Things devices, the volume of data generated at the network edge has grown explosively. As a distributed machine learning solution, Federated Edge Learning (FEL) in mobile edge computing (MEC) provides an effective way to protect privacy by training locally and sharing only model parameters rather than the original data. However, in practical applications, FEL faces two core challenges: First, the data heterogeneity of each edge device can lead to deviations in the global model; The second is how to design an effective incentive mechanism to encourage nodes with limited resources to continuously participate in training. Our aims to simultaneously address these two major challenges by optimizing the local training and global aggregation processes to comprehensively enhance the efficiency and performance of FEL. For this reason, we propose a hybrid model. Firstly, the Stackelberg Stackelberg game model is adopted to describe the relationship between aggregators and edge devices. Meanwhile, the existence of Nash equilibrium is theoretically proved to ensure the stability of the model. Secondly, we propose a novel algorithm named Group Relative Policy Optimization Based on Hierarchical Mean-Field Theory (OGRPO-HMF), which can jointly optimize the local training of nodes and the global model aggregation of servers. We validate the effectiveness and generality of our approach through extensive experimentation on various FEL tasks, showcasing significant performance gains. Extensive experiments on benchmark FEL datasets demonstrate the superior performance of our proposed algorithm, improving the global test accuracy by up to \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$2.26\%$$\end{document} and reducing the global test loss by up to \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$66.27\%$$\end{document} compared to the state-of-the-art counterparts.

Bing Ai, Yu Sun, Jun Wang et al. · 0 citations
Conference Jul 2026

KaaS-Edge: Resource-Aware Knowledge Distillation Service for Heterogeneous Wireless Edge Networks

Federated distillation (FD) enables collaborative edge learning by exchanging soft predictions rather than model parameters, offering communication efficiency and architectural flexibility. However, deploying FD over heterogeneous wireless networks requires principled methods to schedule device participation and allocate upload volumes under per-round resource constraints. Existing approaches assume uniform participation or rely on heuristic selection, ignoring the coupling among communication cost, computational capability, and privacy posture across devices. This paper proposes KaaS-Edge, a Knowledge-as-a-Service framework that formulates device scheduling as budgeted submodular maximization. We derive an optimal water-filling volume allocation in closed form and present RADS (Resource-Aware Distillation Scheduling), a greedy algorithm with a constant-factor approximation guarantee. Experiments on CIFAR-100 demonstrate that KaaS-Edge achieves accuracy comparable to full-participation baselines while reducing per-round communication by nearly ten times and cumulative bandwidth by over an order of magnitude, with graceful degradation under stringent privacy constraints.

Sheng-zhi Huang · 0 citations
Open access Aug 2026

Mobility-Aware and Privacy-Preserving Federated Reinforcement Learning with Multi-Paradigm Machine Learning for Edge Intelligence in 5G/6G Networks

The rise of 5G and 6G networks, along with the rapid growth of edge computing, is creating a strong need for smarter and more privacy-aware ways to handle task offloading as users move across the network. Many current methods still treat mobility prediction, federated learning (FL), and differential privacy (DP) as separate pieces, which often leads to avoidable delays, higher energy use, and weaker data protection. This paper introduces Mobility-Aware Federated Reinforcement Learning (MA-FRL), a framework designed to bring these components together. It integrates deep reinforcement learning (DRL) with supervised and unsupervised ML techniques to enhance edge intelligence, mobility prediction using Markov chains, and Gaussian Differential Privacy (DP) to make better offloading decisions across multi-tier edge environments. MA-FRL uses a federated deep Q-network (DQN), where each edge node trains locally on mobility-aware data and adds DP noise before contributing to the global model. It utilizes NS-3 and م, in addition to real datasets like CRAWDAD, GeoLife, and SPEC power; the framework is among the first to achieve 32% lower latency, 27% energy savings, and strong privacy protection (ε < 1.0). Pareto analysis shows a balance between performance goals and topology-aware tuning, improving results in urban, rural, and vehicular settings. MA-FRL also aligns with the General Data Protection Regulation (GDPR). Future work will explore Long Short-Term Memory (LSTM) and Spatio-Temporal Graph Neural Networks (ST-GNN) mobility models and hardware-in-the-loop testing.

N. Baban · 0 citations
Conference Open access Jul 2026

REVS-T: Trust-Tier-Aware Provider Selection for Secure Vehicular Computation Offloading

Vehicular computation offloading (VCOff) enables resource-constrained vehicles to delegate delay-sensitive tasks to nearby providers. However, it remains vulnerable to strategic malicious nodes that withhold results, or behave intermittently to evade detection. Although reputation values evolve across repeated interactions, long-term security depends on how these signals are governed and enforced at the decision layer. This paper introduces REVS-T, a four-tier governance and tieraware selection mechanism using reputation bands, warningstreak escalation, and pool partitioning. Under persistent attack at 50% adversarial ratio, REVS-T achieves 91.9% task success and 96.7% malicious avoidance with zero false exclusions, outperforming Threshold by 5.5% and Beta by 14.9%. A four-step ablation under shared reputation-update logic shows composite scoring and four-tier governance as the dominant drivers, with ST-conditioned initialization providing phase-shift adaptation and a fairness guarantee no evaluated baseline achieves.

Sharifah Fayi, Ferheen Ayaz, Zhengguo Sheng · 0 citations
Review Open access 2026

Comprehensive Review of Optimization Techniques for User-Centric Distributed Network Slicing in 5G Networks

A QoE-aware framework for Multi-Access Edge Computing-enabled Open Radio Access Network (O-RAN) architectures, combining a graph attention network (GAT) encoder, distributed multi-agent DRL, and privacy-preserving FL, while transitioning control from Quality of Service (QoS) to QoE metrics is proposed.

Manoj Prasad Kunasegran, Wai Leong Pang, S. K. Phang · 0 citations

Related blog posts

Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.