The extent to which an agentic framework can optimize and run an HPC scaling study with a low latency network in Amazon Web Services, accurately transform HPC job specifications between workload managers, and design and run an entire biosciences workflow is assessed.
Abstract
The means to execute and orchestrate software components has changed from human-written code to descriptive prose. In high performance computing, this transition is represented in application orchestration, workload management, and system monitoring and debugging, to name a few. The underlying means to enable descriptive definition of tasks is the use of the Large Language Model with associated tool functions and resources. A combination of a model with access to such resources, modeled in software, encompasses an autonomous framework. As fully automated and agentic frameworks are developed for science, it is important to assess reliability and strategies scoped to specific tasks. In this work, we assess the extent to which an agentic framework can optimize and run an HPC scaling study with a low latency network in Amazon Web Services, accurately transform HPC job specifications between workload managers, and design and run an entire biosciences workflow. We find that the framework completes all three tasks while surfacing task-specific failure modes. In the scaling study, agents deploy and optimize applications but monitor running jobs inefficiently, preferring conservative fixed waits over event subscriptions. In job translation, they convert specifications between Slurm and Flux with high accuracy, with processor-affinity flags the most common error. In the bioscience workflow, the agent reproduces an expert-written variant-calling pipeline almost exactly -- agreeing with the reference call set in 18 of 19 completed runs -- and reaches this result through many distinct yet functionally equivalent workflow implementations. This information is invaluable moving forward to developing multi-cluster setups with scheduling and transformation handled by agents.
RASER is presented, a user-space framework that enables seamless execution of agentic workflows on production HPC clusters by extending Slurm's internal primitives and provides resilience against preemption and failures while maintaining minimal checkpoint/restore overhead.
This paper presents a hierarchical, dynamic architecture and software to discover resources across diverse cloud, edge, and HPC systems and exemplifies the importance of careful coordination between agents, discovery tools, and infrastructure for agentic science.
HPC-AutoResearch is presented, a proof-of-concept system for the autonomous execution of compiled-code research workflows in HPC-like environments that divides this sub-pipeline into five phases—planning, environment setup, coding, compilation, and execution—localizing failures within each phase and enabling iterative...
T. Kotama, Shun-ichiro Hayashi, Daichi Mukunoki et al.· 0 citations
An integrated management framework powered by Tapis is shown how a unified UI-driven workflow streamlines the transition from initial evaluation to deployment, ensuring operational consistency and reproducibility without manual script porting.
Manikya Swathi Vallabhajosyula, Gautam Gururaj Molakalmuru, Samuel Khuvis et al.· Practice and Experience in A...· 0 citations
This paper presents a systematic, practice-driven evaluation of WebAssembly (WASM) as an execution substrate for cloud-native workloads orchestrated through Docker and Kubernetes using runwasi. We develop a reproducible workflow that compiles Rust and Go/TinyGo applications to WASM modules, applies Ahead-of-Time (AOT)...
Álvaro Vázquez-Rodríguez, David Vila-Pérez, Carlos Giraldo-Rodíguez et al.· International Conference on...· 0 citations
Fabrica decouples the API’s structural standards from its application logic, allowing the team to generate working prototypes for immediate testing at Los Alamos National Laboratory while retaining the flexibility to regenerate the service to match pending architectural decisions.
B. McDonald, Alex Lovell-Troy· Practice and Experience in A...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.