A three-layer architecture that addresses the tension between elastic, multi-tenant cloud operations and performance-critical HPC batch workflows, and places cloud-style Infrastructure-as-a-Service (IaaS) management, based on virtualization with GPU passthrough, at the center.
Abstract
The rapid expansion of artificial intelligence (AI) workloads is prompting governments and institutions worldwide to invest in sovereign AI Factories: large-scale infrastructures designed to support both cloud-native AI services and AI-oriented HPC workloads. This paper proposes a three-layer architecture that addresses the tension between elastic, multi-tenant cloud operations and performance-critical HPC batch workflows. We position the Cloud Layer as the central orchestration and resource management plane, responsible for provisioning GPU-accelerated virtual machines with near-bare-metal throughput and for coordinating compute, GPU, network, and storage resources under multi-tenant isolation. Unlike Kubernetes-centric designs, our architecture places cloud-style Infrastructure-as-a-Service (IaaS) management, based on virtualization with GPU passthrough, at the center. This provides strong multi-tenancy, clearer security boundaries, and hardware-level isolation between tenants, while remaining agnostic to the workload orchestrator. We ground the proposal in a concrete reference implementation: the OpenNebula AI Factory Reference Architecture, which is used in large-scale European initiatives such as IPCEI-CIS. Experimental evaluation with the cuBLAS benchmark confirms that GPU passthrough virtualization introduces negligible performance overhead compared with bare metal across three representative GPU generations (NVIDIA GB200, H100, and L40S), including configurations using MIG partitioning, for compute-bound workloads. We analyze key challenges in resource allocation, placement, and partitioning, and discuss design trade-offs and open research directions for efficient, sovereign AI infrastructures.
This article presents a practitioner-oriented engineering framework for provisioning artificial intelligence-ready infrastructure that remains architecturally stable across accelerator generations, managed service evolutions, and organizational growth trajectories, and projects the long-term strategic implications of m...
Hemanth Kumar Gandavarapu· International Journal of Eng...· 0 citations
The key message is that there is a new central design coordination of the cloud architecture that enables sustainable growth of the cloud, not merely increasing incremental capacity on hardware.
Priyadarshni Shanmugavadivelu· International journal of com...· 0 citations
This paper explores hybrid cloud-edge infrastructures as a scalable solution for deploying AI in IIoT environments and presents an architectural framework that balances compute-intensive model training in the cloud with low-latency inference at the edge.
Jennifer Clark· International Journal of Mac...· 0 citations
The paper argues that multi-tenancy must be treated not merely as a virtualization problem but as a data-governance, workload-management, and responsible-AI problem, and provides a foundation for scalable AI data lakes while identifying limitations related to resource interference, governance complexity, data heterogen...
Arjun Mehta, Priya Sharma· International Journal of Adv...· 0 citations
An integrated management framework powered by Tapis is shown how a unified UI-driven workflow streamlines the transition from initial evaluation to deployment, ensuring operational consistency and reproducibility without manual script porting.
Manikya Swathi Vallabhajosyula, Gautam Gururaj Molakalmuru, Samuel Khuvis et al.· Practice and Experience in A...· 0 citations
An architecture that decouples complex policy enforcement from high-speed packet forwarding to support VPC semantics on back-end NICs and enable front-end/back-end integration is proposed, suggesting that commodity hardware can support both high-throughput AI training and flexible VPC features.
Yinhe Wang, Xing Li, Enge Song et al.· Asia-Pacific Workshop on Net...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.