Skip to content

Hybrid Workflow Composition for Extreme-Scale Data Processing: A Case Study on the HL-LHC (Extended Version)

Jul 2026 · arXiv.org · Vol abs/2607.26877 · 0 citations · 20 references
Computer Science

TL;DR

This paper presents a novel simulation framework for characterizing the interplay between taskset granularity and system-level constraints and demonstrates that hybrid composition strategies, which dynamically balance taskset independence with execution grouping, can yield up to 3.8x throughput increase and a 14.9x reduction in network overhead.

Abstract

High-Throughput Computing (HTC) environments tailored for high-concurrency resource efficiency require sophisticated orchestration to manage petabyte-scale data across heterogeneous resources. A critical but often overlooked challenge is workflow composition: the strategic grouping of tasksets within a Directed Acyclic Graph (DAG) to mitigate execution overhead while maximizing resource utilization. This paper presents a novel simulation framework for characterizing the interplay between taskset granularity and system-level constraints (e.g., job latency, failure rate, throughput, and I/O bandwidth). By exploring a high-dimensional parameter space, we quantify the performance sensitivity of diverse workflow topologies. Our results demonstrate that hybrid composition strategies, which dynamically balance taskset independence with execution grouping, can yield up to 3.8x throughput increase and a 14.9x reduction in network overhead. We further propose a multi-metric objective function that enables policy-driven optimization, allowing system architects to navigate the Pareto frontier between throughput, I/O cost, and CPU efficiency. These findings provide a rigorous foundation for automated workflow synthesis in distributed systems, offering a scalable model for next-generation scientific pipelines. All artifacts are publicly available.

View source

Similar papers

Open access Sep 2026

A Structure-Aware Hybrid Scheduling Framework for Mixed-Dependency Workflow Scheduling in V2X Testing

This paper presents AOE–CP (AON DAG with Edge-Weighted Transformation and Critical Path Scheduling), a structure-aware hybrid scheduling architecture for V2X testing that achieves performance gains through domain-specific structural reorganization rather than new scheduling rules.

Zhu-Hua Zhang, Ning Ye, Chong-Yang Wang et al. · 0 citations
Open access 2024

Intelligent Workflow Scheduling for Distributed Data Processing Systems

An intelligent workflow scheduling framework that improves performance through adaptive decision-making, predictive analytics, and machine learning, and addresses key challenges like load balancing, scalability, energy efficiency, and fault tolerance is proposed.

D. Parnas · 0 citations
Review Open access Jul 2026

Enhancing the Kubernetes Scheduler: A State-of-the-Art Review from Cloud to Edge

A comprehensive review of Kubernetes scheduling strategies published between January 2023 and January 2026 is presented and a multi-dimensional taxonomy is established that categorizes scheduling approaches based on common objectives, modification methods, optimization methodologies, targeted workloads, evaluation meth...

Mohammed Alhakimi, R. Latip · 0 citations
Aug 2026

A Unified Bandwidth Orchestration Framework for Hierarchical Data Storage Systems

This paper presents an in-depth analysis of data migration behavior in commercial HSSs, uncovering substantial performance variability when multiple migration tasks execute concurrently, and proposes PASCAL, a system-level bandwidth orchestration framework that improves performance robustness in production-grade HSSs.

Ji Zhang, Li Liu, André Brinkmann et al. · 0 citations
Book Open access Aug 2026

G-STAR: Graph-based Scheduling with Trace-driven Adaptive Routing for Industrial LLM-based Multi-Agent Systems

Large Language Model-based Multi-Agent Systems (LLM-MAS) have shown exceptional promise for complex tasks, including retrieval-augmented generation and autonomous data analytics. However, their deployment in resource-constrained industrial environments faces critical challenges, such as unpredictable end-to-end latency...

Jia-Bao Song, Yun-Sheng Xia, Bei-Bei Kong et al. · 0 citations
2026

Workflow-Aware Expert Routing for Distributed LLM Serving Over the Edge-Cloud Continuum

Deploying Large Language Models (LLMs) over the edge-cloud continuum faces severe stability challenges due to the conflict between stochastic network topology and complex workflow dependencies. Existing schedulers, relying either on computationally prohibitive Graph Neural Networks (GNNs) or topology-agnostic heuristic...

Yan Gao, Shaoyuan Huang, Yonghui Ye et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.