Scalable AI Inference Pipelines Across Edge and Cloud Computing Environments
Abstract
The increasing deployment of artificial intelligence (AI) applications in healthcare, industrial Internet of Things (IIoT), intelligent transportation, and next-generation wireless systems has created a demand for inference architectures that simultaneously provide low latency, scalability, privacy, reliability, and efficient resource utilization. Conventional cloud-centric inference architectures provide substantial computational capacity but can introduce network latency, bandwidth consumption, privacy exposure, and dependence on centralized infrastructure. Edge computing addresses several of these limitations by relocating computation closer to data sources, while cloud environments remain important for computationally intensive and globally coordinated workloads. This research examines a scalable edge-to-cloud AI inference pipeline in which inference tasks are dynamically distributed across heterogeneous edge and cloud resources. The methodology synthesizes the provided literature on federated learning, edge resource allocation, dynamic scheduling, privacy preservation, machine learning for 6G, IIoT, and secure healthcare systems. A layered architectural model is developed around workload characterization, adaptive task placement, communication-aware scheduling, privacy protection, and resilient orchestration. The analysis indicates that scalability is not achieved merely by adding computational resources; rather, it depends on coordinated optimization of computation, communication, privacy, and scheduling. The proposed conceptual framework positions edge inference as the first computational layer, cloud inference as an elastic computational layer, and intelligent orchestration as the mechanism connecting the two. The resulting architecture provides a basis for resilient real-time AI systems while highlighting unresolved challenges involving heterogeneous hardware, dynamic workloads, privacy-utility trade-offs, and cross-layer optimization.