Skip to content
Review Open access

Towards a Unified Framework for Large–Small Model Collaboration in Cloud–Edge Systems

Jul 2026 · Journal of Artificial Intelligence for Automation · Vol 1 · 0 citations · 125 references

TL;DR

This survey focuses on how cloud-side large models and edge-side small models share, update, and coordinate knowledge and identifies open challenges for trustworthy, sustainable, and self-evolving cloud–edge collaborative intelligence.

Abstract

Foundation models have improved the reasoning and generation ability of artificial intelligence systems. However, they are difficult to deploy in edge environments with limited computation, memory, and data access. Small models are easier to run on edge devices. They support fast and low-latency inference, but they often lack global semantic reasoning and cross-domain generalization. This gap between model ability and deployment cost motivates large–small model collaboration in cloud–edge systems. This survey provides a systematic review and a knowledge-floworiented taxonomy of such collaboration. It focuses on how cloud-side large models and edge-side small models share, update, and coordinate knowledge. We review knowledge distillation, split inference, federated and continual adaptation, and elastic offloading. We also cover lightweight deployment, modular expert design, privacyaware coordination, and agent-driven orchestration. Unlike surveys on edge intelligence, federated learning, model compression, TinyML, or cloud–edge resource scheduling, this survey centers on model collaboration. We treat large–small collaboration as a knowledge-centered problem linked to real deployment constraints. We further discuss trade-offs in accuracy, latency, bandwidth, privacy, energy efficiency, adaptability, and lifecycle management. Finally, we identify open challenges for trustworthy, sustainable, and self-evolving cloud–edge collaborative intelligence.

Read PDF

Similar papers

Jul 2026

A Cloud Continuum Research Infrastructure for Distributed CPS Experimentation

The proposed approach separates the research-infrastructure layer, which exposes and manages distributed resources, from the application layer, where Cyber-Physical workflows are organized according to an Edge-Fog-Cloud pattern in which placement, timing, and data provenance are treated as first-class experimental conc...

Fabio Orazio Mirto, Giuseppe Tricomi, L. D’Agati et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Smart Adaptive Computing Across the Continuum: LLMs in IoT-Edge-Cloud Resource Management

Managing resources across IoT, edge, and cloud layers calls for continuous, context-aware decisions under constraints that rarely stay fixed. Deep reinforcement learning (DRL) handles this class of problems well, and large language models (LLMs) are increasingly used to augment DRL pipelines, yet the architectural rela...

Antonino Vaccarella, Lan-Pei Li, Vincenzo Lomonaco et al. · 0 citations
Aug 2026

From Cloud to Crowd: Democratizing LLM Service With Decentralized Edge Collaboration for RAG

Results show that DEFRAG narrows the SLM-LLM accuracy gap, while reducing cost by up to 98.4% and increasing peak throughput by up to 97.8% over centralized services, demonstrating the potential of DEFRAG for democratized LLM services at the edge.

Jiaxing Li, Heng-Zhi Wang, Feng Wang et al. · 0 citations
Open access Sep 2026

FinDS-Agent: A Cloud–Edge Collaborative Data Science Agent for Financial Analytics

Large language model agents can automate data science workflows, but cloud-centric deployment exposes sensitive context and edge-only deployment limits analytical capability. We present FinDS-Agent, a cloud–edge framework that keeps raw records and program execution at the trusted edge while providing a policy-screened...

Xiao-Zheng Du, Rui-Jun Deng, Cheng Wang et al. · 0 citations
#machine learning Preprint Sep 2026

HoliBench: A Cross-Platform Benchmarking and Deployment Toolkit for Foundation Models in CPS-IoT Applications

Foundation models, including large language models, vision-language models, and time-series foundation models, are increasingly deployed on embedded and edge platforms for CPS and IoT applications, where energy, latency, and memory are as critical as task accuracy. Existing benchmarking tools evaluate model capability...

Inesh Chakrabarti, Ze-Jun Xiong, Pragya Sharma et al. · 0 citations
#graph neural networks Open access Sep 2026

Synergizing large and small models in cloud-edge continuum: a spatiotemporal hypergraph approach for dynamic offloading

This paper proposes a spatiotemporal hypergraph-driven framework integrating high-order topological feature extraction with dynamic resource modeling, and introduces dynamic hypergraph sequences to naturally encompass local conflict domains, mitigating the topological blind spots and “over-smoothing” issues inherent in...

Kun Ding, Xi-Wen Qiu, Nian-Feng Weng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.