EASE-CloudNet is a two-phase safety-alignment framework for generative small language models (SLMs) deployed on resource-constrained edge nodes that formulate deployment cost as a differentiable gate-conditioned expectation, so measured latency and energy constants affect the router through its reasoning probability.
Abstract
EASE-CloudNet is a two-phase safety-alignment framework for generative small language models (SLMs) deployed on resource-constrained edge nodes. Its input is a natural-language user query and its output is a safe, helpful natural-language response or refusal; network-traffic classification and resource-scheduling actions are outside the task evaluated in this study. In Phase 1, a cloud teacher uses a security policy graph to generate structured safety rationales and response targets, which are distilled into Qwen2.5-1.5B/3B and Llama3.2-3B students. In Phase 2, an offline heterogeneous graph and a two-layer GraphSAGE model identify vulnerable semantic regions; these vulnerability targets supervise a lightweight edge-side router. We formulate deployment cost as a differentiable gate-conditioned expectation, so measured latency and energy constants affect the router through its reasoning probability. In the Qwen2.5-1.5B ablation experiments, the full model obtains 3.9% StrongREJECT ASR, 54.7% MMLU accuracy, and 70 average generated tokens; an A100 reference profile reports 18.8 ms/query and 2.37 J/query, or 3.3% latency and 2.6% measured GPU-energy overhead over the unaligned model. Physical edge runs measured a direct/reasoning end-to-end latency of 41.2/68.7 ms on Jetson Orin NX and 62.5/105.3 ms on Snapdragon 8 Gen 3, with a direct/reasoning energy of 0.48/0.79 and 0.71/1.18 J/query, respectively. Equal-seed Holm–Bonferroni-corrected tests confirm lower ASR than EASE on StrongREJECT and WildJailbreak for all three base models (p<0.01).
Network slicing in multidomain 5G environments introduces critical challenges in guaranteeing Service Level Agreements (SLAs) under dynamic resource variability and heterogeneous service requirements. This article presents TriSLA, a closed-loop, preventive, SLA-aware architecture designed to evaluate feasibility at req...
Abel J. R. Lisboa, G. Z. Bruno, C. B. Both· 0 citations
FinDS-Agent is presented, a cloud–edge framework that keeps raw records and program execution at the trusted edge while providing a policy-screened, sanitized context to support cloud planning and support selective cloud planning while delimiting statistical, privacy, and transfer claims.
Xiao-Zheng Du, Rui-Jun Deng, Cheng Wang et al.· Future Internet· 0 citations
Cloud–edge collaborative inference has emerged as a promising paradigm to address the latency, energy, and privacy challenges of large language models (LLMs). However, current offloading mechanisms often struggle to efficiently capture the dynamic semantic dependencies between historical context and ongoing queries. Th...
Xianzhong Tian, Yu Wang, Guanpeng Zhu· IEEE Internet of Things Jour...· 0 citations
This work presents Sentinel-RL, an agentic-SOC architecture that decouples topological reasoning from semantic reasoning, and contributes a reusable engineering pattern, a portable HPC deployment pattern, and an enterprise-readiness analysis covering false-positive economics, reversibility guarantees, audit compliance,...
Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild· 1 citation
G-STAR is a general graph-based scheduling framework that formalizes complex MAS pipelines as attributed Directed Acyclic Graphs (DAGs) and develops an industry-grade orchestration stack with asynchronous execution, resilient serving, and audit-friendly artifacts, offering a practical solution for optimizing web-scale...
Jia-Bao Song, Yun-Sheng Xia, Bei-Bei Kong et al.· Proceedings of the 32nd ACM...· 0 citations