Nix's consistent package layout, full environment isolation, and flake-based composition resolve the dependency discovery, leakage, and composition problems the authors encountered, while unifying C/C++ and Python management under a single declarative specification that also generates the deployment container.
Abstract
Reproducibility in HPC remains difficult under the constraints of production supercomputers: no root access, limited internet, and software stacks that increasingly span C/C++, Fortran, Python, MPI, and GPU runtimes. Traditional approaches based on environment modules and Conda require manual intervention to locate dependencies, leak system libraries into builds, and fail to compose across projects. Containers help with deployment but do not by themselves guarantee reproducibility. We report on our experience building a hybrid HPC/AI software stack with Nix, covering local development on a workstation without root and remote deployment as an Apptainer image on a production cluster. Nix's consistent package layout, full environment isolation, and flake-based composition resolve the dependency discovery, leakage, and composition problems we encountered, while unifying C/C++ and Python management under a single declarative specification that also generates the deployment container. We discuss trade-offs against Spack and Guix, the development-versus-production split addressed via CMake presets, and current gaps in ML package coverage in Nixpkgs.
(English) High-performance computing (HPC) platforms are evolving towards increasingly complex architectures: many-core CPUs with multi-level NUMA hierarchies, heterogeneity with multiple classes of accelerators and higher-capacity interconnects. The increasing complexity and variety of resources in these machines make...
HPC-AutoResearch is presented, a proof-of-concept system for the autonomous execution of compiled-code research workflows in HPC-like environments that divides this sub-pipeline into five phases—planning, environment setup, coding, compilation, and execution—localizing failures within each phase and enabling iterative...
T. Kotama, Shun-ichiro Hayashi, Daichi Mukunoki et al.· 0 citations
Coding agents have become real users of high-performance computing (HPC) systems, yet today's HPC abstractions, interfaces, and policies remain designed for human-driven workflows. In our measurement, users running coding agents are only 19.5% of the observed population, but account for 55.8% of job submissions, 29.1%...
Yun-Jia Zheng, Bintang Dwi Marthen, Zachary Pan et al.· 0 citations
We propose PKDB, the first interactive debugger for GPU and multithreaded low-level kernels written in Python. Python is widely used in high performance computing (HPC), with frameworks such as PyKokkos translating Python-embedded domain-specific languages to native code that runs across OpenMP-threaded CPUs and variou...
Ivan Grigorik, Gabriel Kosmacher, G. Biros et al.· 0 citations
Low code development platforms like OutSystems allow teams to build data-centric applications using visual models that the platform takes care of the database plumbing [1] [2]. We have demonstrated in an earlier study that the write path is row-wise and fails at scale; and in a previous study we introduced EAVS [25], a...
Anamika Garg, Sachin H. Patel· International journal of com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.