Skip to content
Conference

Data-Parallel and Heterogeneity-Aware Scheduling for Geo-Distributed Serverless Scientific Workflows

Jul 2026 · IEEE International Conference on Cloud Computing · pp. 78-88 · 0 citations · 36 references

Abstract

Serverless computing has emerged as a promising paradigm for executing scientific workflows characterized by complex task dependencies, data-intensive operations, and high computational demands. However, most existing scheduling approaches assume homogeneous, single-region environments and primarily focus on isolated function execution. These assumptions overlook two critical challenges: (i) the inherent data-parallel nature of workflow tasks, and (ii) the heterogeneity of computing resources across geo-distributed serverless platforms. In this paper, we address these limitations by proposing a novel scheduling framework for geo-distributed serverless environments that explicitly models intra-function data parallelism, heterogeneous abstract resources, and regional concurrency constraints. We formulate a makespan minimization problem in which each function can either execute entirely on a single high-capacity resource or be partitioned across multiple heterogeneous lower-capacity resources, subject to region-specific concurrency limits.To solve this problem, we design a Deep Q-Network (DQN)-based scheduler augmented with two auxiliary heuristics. The first, Critical Workload First, prioritizes high-workload functions through an exhaustive split-deployment search over heterogeneous abstract serverless resources. The second, Load-Aware Heuristic, selects execution regions using a weighted load metric combined with penalty-based resource assignment. We evaluate our approach on five representative scientific workflows BWA, Montage, Inspiral, CyberShake, and SIPHT and using real-world round-trip time measurements from Azure Function deployments across three geo-distributed regions. Experimental results demonstrate that our DQN-based scheduler reduces makespan by up to 29.13% compared to state-of-the-art approaches.

View source

Similar papers

Conference Jul 2026

AMSche: Affinity-Aware Microservice Scheduling for Communication-Intensive Tasks

As computing resources in cloud environments become increasingly abundant, executing complex scientific workflows on large-scale cloud infrastructure has become a standard practice. However, communication-intensive workflows face two fundamental bottlenecks. First, the lack of physical topology awareness often forces h...

Hao Wei, Hai-Liang Chen, Jia-Nan Sun et al. · 0 citations
Conference Jul 2026

WaSMa: Workflow-Aware Scheduling for Multi-Agent LLM Systems

Public LLM services serve diverse multi-agent applications with varying workflow dependencies and performance requirements. Requests generated by these applications often exhibit commonality and interdependence, yet current systems largely ignore such application-level structure. As a result, at the LLM engine cluster...

Uttam Rao, Ali Zafar Sadiq, Hai-Ying Shen et al. · 0 citations
Jul 2026

Hybrid Workflow Composition for Extreme-Scale Data Processing: A Case Study on the HL-LHC (Extended Version)

This paper presents a novel simulation framework for characterizing the interplay between taskset granularity and system-level constraints and demonstrates that hybrid composition strategies, which dynamically balance taskset independence with execution grouping, can yield up to 3.8x throughput increase and a 14.9x red...

A. M. Rodrigues, D. Thain · 0 citations
Preprint Sep 2026

Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs

Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability. The logical workflow defines the required computation, whereas its physical scheduling units, model-lifecycle actions, r...

Jing-Hao Wang, Yi-Feng Zhang, Xiao Zhou et al. · 0 citations
Open access Aug 2026

Computing Resource-Aware Operation Optimization Strategy for MPI Jobs in Cloud-Native Environment

A resource-aware optimization framework that dynamically selects the MPI process count and performs node- and NUMA-aware process placement and reduces task-sequence execution time and improves the evaluated resource-utilization metrics by more than 30%.

Wenxiao Wang, Zi-Bo Gao, Guoding Ji et al. · 0 citations
Open access 2026

Adaptive Spotted Hyena Optimizer for Latency-Aware Task Scheduling in Heterogeneous Multicore Systems

The Latency-Aware Adaptive Spotted Hyena Optimizer (LA-ASHO) is proposed, a novel metaheuristic scheduling framework grounded in the social hunting behaviour of spotted hyenas that achieves statistically significant reductions in workflow completion latency relative to established baseline schedulers such as; Min-Min,...

Igiri C. G, Ejekwu Obunezi, Ujah Alechenu Israel · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.