Skip to content

Performance evaluation of scheduling tasks in many-core systems utilizing processes and threads

Jul 2026 · arXiv.org · Vol abs/2607.04821 · 0 citations · 29 references
Computer Science

TL;DR

Lightweight thread scheduling is optimal for shared-memory row sorting, while AIMD/adaptive scheduling and pipe-based process scheduling remain valuable for contention-aware execution, explicit inter-process coordination, and distributed-style heterogeneous workload management.

Abstract

This study assesses the scalability of process-based and thread-based schedulers for many-core shared-memory systems using a memory-intensive row-wise quick-sort workload on large three-dimensional tensors. The process-based evaluation considers bounded prolific, bounded collective, and three pipe-based producer-consumer schedulers: one-to-one, one-to-many, and many-to-many. These pipe schedulers dynamically stream task identifiers to worker processes, exchanging increased inter-process communication overhead for enhanced runtime load balancing and flexible chunk-based task dispatching. The thread-based evaluation examines static, dynamic, guided, chunk-based, chunk-stealing, adaptive chunk, and AIMD adaptive scheduling strategies. The AIMD scheduler employs an additive-increase multiplicative-decrease policy inspired by TCP congestion control, utilizing an exponentially weighted moving average (EWMA) of CPU utilization to regulate a contention window that limits the number of concurrently active chunks. The adaptive chunk scheduler further modifies chunk size based on observed per-thread execution speed. Experimental results on a 24-core x86-64 platform indicate that thread schedulers deliver the highest overall performance, with dynamic and guided scheduling yielding the most favorable practical outcomes. Among process schedulers, pipe-based designs demonstrate the strongest scalability, with one-to-one pipes excelling for smaller workloads and many-to-many pipes preferred for larger workloads. In summary, lightweight thread scheduling is optimal for shared-memory row sorting, while AIMD/adaptive scheduling and pipe-based process scheduling remain valuable for contention-aware execution, explicit inter-process coordination, and distributed-style heterogeneous workload management.

View source

Similar papers

Future Generation Computer Systems

Concord, a novel GPU sharing-enabled workload scheduler that outperforms state-of-the-art schedulers, achieves a 1.68 × reduction in JCT and a 29% improvement in GPU utilization in high-load scenarios.

Xin-Hua Wang, Wei-Wei Lin, Hai-Jie Wu et al. · 0 citations
Open access Aug 2026

Kernel-Level Dynamic Priority Scheduling for Containers

A dynamic priority scheduling framework at the kernel level that enhances the CPU allocation to latency-sensitive containers running in Kubernetes environments and reveals a significant improvement in terms of latency reduction, enhanced throughput, efficient utilization of CPU resources, and stable performance of sche...

T. Rajkumar, Nishanth D., P. M et al. · 0 citations
Open access Aug 2026

Computing Resource-Aware Operation Optimization Strategy for MPI Jobs in Cloud-Native Environment

A resource-aware optimization framework that dynamically selects the MPI process count and performs node- and NUMA-aware process placement and reduces task-sequence execution time and improves the evaluated resource-utilization metrics by more than 30%.

Wenxiao Wang, Zi-Bo Gao, Guoding Ji et al. · 0 citations
Preprint Aug 2026

A Smallest-Need-First Job Scheduling Framework with Adaptive Optimization of Idle Node Counts for Energy-Efficient HPC Systems

Power-state management in high-performance computing (HPC) clusters must reduce idle energy without excessive wake-up delays for rigid parallel jobs. This paper presents SNF-ICON, an event-driven controller combining smallest-need-first (SNF) gang scheduling, predictive wake timing, and adaptive warm-spare control. At...

Reza Pulungan, Raka Satya Prasasta, Santana Yuda Pradata et al. · 0 citations
Jul 2026

CW-Ghost: Search-Free Granularity Selection for Helper-Thread Prefetching via Capacity Windows

Helper-thread prefetching hides the latency of irregular memory accesses by executing address dependency chains ahead of the main thread. However, its effectiveness depends on the range of future iterations covered by the helper thread. A fixed coverage range cannot consistently accommodate different workloads and proc...

Ya Zhang, Tong Lei, Yao Chen et al. · 0 citations
Jul 2026

CrocSort: Resource-Efficient, Skew-Resilient Parallel External Merge Sort

CrocSort is presented, a byte-balanced parallel external merge sort with configurable memory and per-phase thread settings with practical resource-configuration rules for selecting these settings from input size, memory budget, and thread cap.

Riki Otaki, Charles Benello, Fuheng Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.