Jul 2026· IEEE International Conference on Cloud Computing· pp. 78-88· 0 citations· 36 references
Abstract
Serverless computing has emerged as a promising paradigm for executing scientific workflows characterized by complex task dependencies, data-intensive operations, and high computational demands. However, most existing scheduling approaches assume homogeneous, single-region environments and primarily focus on isolated function execution. These assumptions overlook two critical challenges: (i) the inherent data-parallel nature of workflow tasks, and (ii) the heterogeneity of computing resources across geo-distributed serverless platforms. In this paper, we address these limitations by proposing a novel scheduling framework for geo-distributed serverless environments that explicitly models intra-function data parallelism, heterogeneous abstract resources, and regional concurrency constraints. We formulate a makespan minimization problem in which each function can either execute entirely on a single high-capacity resource or be partitioned across multiple heterogeneous lower-capacity resources, subject to region-specific concurrency limits.To solve this problem, we design a Deep Q-Network (DQN)-based scheduler augmented with two auxiliary heuristics. The first, Critical Workload First, prioritizes high-workload functions through an exhaustive split-deployment search over heterogeneous abstract serverless resources. The second, Load-Aware Heuristic, selects execution regions using a weighted load metric combined with penalty-based resource assignment. We evaluate our approach on five representative scientific workflows BWA, Montage, Inspiral, CyberShake, and SIPHT and using real-world round-trip time measurements from Azure Function deployments across three geo-distributed regions. Experimental results demonstrate that our DQN-based scheduler reduces makespan by up to 29.13% compared to state-of-the-art approaches.
As computing resources in cloud environments become increasingly abundant, executing complex scientific workflows on large-scale cloud infrastructure has become a standard practice. However, communication-intensive workflows face two fundamental bottlenecks. First, the lack of physical topology awareness often forces h...
Hao Wei, Hai-Liang Chen, Jia-Nan Sun et al.· Fall Joint Computer Conferen...· 0 citations
Public LLM services serve diverse multi-agent applications with varying workflow dependencies and performance requirements. Requests generated by these applications often exhibit commonality and interdependence, yet current systems largely ignore such application-level structure. As a result, at the LLM engine cluster...
Uttam Rao, Ali Zafar Sadiq, Hai-Ying Shen et al.· International Conference on...· 0 citations
This paper presents a novel simulation framework for characterizing the interplay between taskset granularity and system-level constraints and demonstrates that hybrid composition strategies, which dynamically balance taskset independence with execution grouping, can yield up to 3.8x throughput increase and a 14.9x red...
Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability. The logical workflow defines the required computation, whereas its physical scheduling units, model-lifecycle actions, r...
Jing-Hao Wang, Yi-Feng Zhang, Xiao Zhou et al.· 0 citations
A resource-aware optimization framework that dynamically selects the MPI process count and performs node- and NUMA-aware process placement and reduces task-sequence execution time and improves the evaluated resource-utilization metrics by more than 30%.
Wenxiao Wang, Zi-Bo Gao, Guoding Ji et al.· Journal of Intelligent Compu...· 0 citations
The Latency-Aware Adaptive Spotted Hyena Optimizer (LA-ASHO) is proposed, a novel metaheuristic scheduling framework grounded in the social hunting behaviour of spotted hyenas that achieves statistically significant reductions in workflow completion latency relative to established baseline schedulers such as; Min-Min,...
Igiri C. G, Ejekwu Obunezi, Ujah Alechenu Israel· Journal of Artificial Intell...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.