Skip to content
Preprint

RASER: Resilient Agent Scheduling and Execution Runtime for HPC Clusters

Sep 2026 · 1 citation · 17 references
Computer Science

TL;DR

RASER is presented, a user-space framework that enables seamless execution of agentic workflows on production HPC clusters by extending Slurm's internal primitives and provides resilience against preemption and failures while maintaining minimal checkpoint/restore overhead.

Abstract

The emergence of modern agents powered by large language models has created a demand for executing long-horizon, autonomous workflows in various domains that require significant computational resources. While High Performance Computing clusters provide the ideal infrastructure for these computation-intensive workloads, traditional HPC job schedulers such as Slurm are not designed for dynamic, agentic workflows characterized by unpredictable task durations, external API calls, and fault tolerance requirements of modern agents. This work presents RASER, a user-space framework that enables seamless execution of agentic workflows on production HPC clusters by extending Slurm's internal primitives. RASER introduces agentic job arrays with work stealing via shared filesystem queues, user-space checkpointing through application-level state serialization combined with Slurm requeue, and Apptainer container-based isolation without requiring any image modifications. Evaluations demonstrate that RASER reduces makespan by nearly 39% compared to static partitioning while achieving near-full CPU utilization. RASER provides resilience against preemption and failures while maintaining minimal checkpoint/restore overhead. It requires no kernel privileges or external database infrastructure, making it an accessible solution for deploying agentic workflows on existing HPC infrastructure.

View source

Similar papers

Preprint Sep 2026

Towards Efficient HPC Systems for Agents: Challenges and Opportunities

Coding agents have become real users of high-performance computing (HPC) systems, yet today's HPC abstractions, interfaces, and policies remain designed for human-driven workflows. In our measurement, users running coding agents are only 19.5% of the observed population, but account for 55.8% of job submissions, 29.1%...

Yun-Jia Zheng, Bintang Dwi Marthen, Zachary Pan et al. · 0 citations
Preprint Sep 2026

XMPIaaS: Towards Cloud Native MPI via Cooperative Process Migration

Message Passing Interface (MPI) has been the dominant programming model for High Performance Computing (HPC) for three decades, and as HPC workloads increasingly migrate to cloud infrastructure for scalability and cost efficiency, MPI applications must contend with an execution environment fundamentally unlike traditio...

Shun-Yu Yao, Dimitrios S. Nikolopoulos, A. R. Butt · 0 citations
Preprint Sep 2026

Performance vs Portability in Heterogeneous HPC Environments: Why Pre-execution Benchmarking is Required

Cloud computing and high-performance computing (HPC) typically follow different paradigms: cloud services are often orchestrated using Kubernetes, whereas HPC workloads are managed through batch schedulers such as Slurm. Growing demand for shared computational resources increases the need for interoperability between t...

M. Mačernis · 0 citations
#artificial intelligence Preprint Sep 2026

ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations

Traditional scientific computing requires researchers to translate intent into environment configuration, resource requests, and executable jobs, then diagnose failures from scheduler state and logs. We present ASCEND (Autonomous Scientific Computing Engine and Novel Discovery), an AI-powered agent interface that suppo...

J. P. Liu, Uthpala Herath, Andrew Petersen · 0 citations
Preprint Aug 2026

Hierarchical Server Architecture for Agentic Science

This paper presents a hierarchical, dynamic architecture and software to discover resources across diverse cloud, edge, and HPC systems and exemplifies the importance of careful coordination between agents, discovery tools, and infrastructure for agentic science.

Vanessa V. Sochat, Daniel Milroy · 1 citation
Open access Aug 2026

Computing Resource-Aware Operation Optimization Strategy for MPI Jobs in Cloud-Native Environment

A resource-aware optimization framework that dynamically selects the MPI process count and performs node- and NUMA-aware process placement and reduces task-sequence execution time and improves the evaluated resource-utilization metrics by more than 30%.

Wenxiao Wang, Zi-Bo Gao, Guoding Ji et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.