Skip to content
Open access

A Layered Architecture for AI Factories

2026 · IEEE Access · Vol 14, pp. 104171-104187 · 0 citations · 85 references
Computer Science

TL;DR

A three-layer architecture that addresses the tension between elastic, multi-tenant cloud operations and performance-critical HPC batch workflows, and places cloud-style Infrastructure-as-a-Service (IaaS) management, based on virtualization with GPU passthrough, at the center.

Abstract

The rapid expansion of artificial intelligence (AI) workloads is prompting governments and institutions worldwide to invest in sovereign AI Factories: large-scale infrastructures designed to support both cloud-native AI services and AI-oriented HPC workloads. This paper proposes a three-layer architecture that addresses the tension between elastic, multi-tenant cloud operations and performance-critical HPC batch workflows. We position the Cloud Layer as the central orchestration and resource management plane, responsible for provisioning GPU-accelerated virtual machines with near-bare-metal throughput and for coordinating compute, GPU, network, and storage resources under multi-tenant isolation. Unlike Kubernetes-centric designs, our architecture places cloud-style Infrastructure-as-a-Service (IaaS) management, based on virtualization with GPU passthrough, at the center. This provides strong multi-tenancy, clearer security boundaries, and hardware-level isolation between tenants, while remaining agnostic to the workload orchestrator. We ground the proposal in a concrete reference implementation: the OpenNebula AI Factory Reference Architecture, which is used in large-scale European initiatives such as IPCEI-CIS. Experimental evaluation with the cuBLAS benchmark confirms that GPU passthrough virtualization introduces negligible performance overhead compared with bare metal across three representative GPU generations (NVIDIA GB200, H100, and L40S), including configurations using MIG partitioning, for compute-bound workloads. We analyze key challenges in resource allocation, placement, and partitioning, and discuss design trade-offs and open research directions for efficient, sovereign AI infrastructures.

Read PDF

Similar papers

Open access Jul 2026

Provisioning AI-Ready Infrastructure at Scale: Engineering Considerations for Infrastructure Architects

This article presents a practitioner-oriented engineering framework for provisioning artificial intelligence-ready infrastructure that remains architecturally stable across accelerator generations, managed service evolutions, and organizational growth trajectories, and projects the long-term strategic implications of m...

Hemanth Kumar Gandavarapu · 0 citations
Open access 2020

Hybrid Cloud-Edge Infrastructures for Scalable IIoT AI Deployments

This paper explores hybrid cloud-edge infrastructures as a scalable solution for deploying AI in IIoT environments and presents an architectural framework that balances compute-intensive model training in the cloud with low-latency inference at the edge.

Jennifer Clark · 0 citations
Open access Aug 2026

Multi-Tenant Data Lake Architecture for Scalable AI and Big Data Workload Management

The paper argues that multi-tenancy must be treated not merely as a virtualization problem but as a data-governance, workload-management, and responsible-AI problem, and provides a foundation for scalable AI data lakes while identifying limitations related to resource interference, governance complexity, data heterogen...

Arjun Mehta, Priya Sharma · 0 citations
Book Open access Jul 2026

From HPC to Edge: A Web-Based Workflow for AI Model Testing and Deployment

An integrated management framework powered by Tapis is shown how a unified UI-driven workflow streamlines the transition from initial evaluation to deployment, ensuring operational consistency and reproducibility without manual script porting.

Manikya Swathi Vallabhajosyula, Gautam Gururaj Molakalmuru, Samuel Khuvis et al. · 0 citations
Book Open access Aug 2026

Integrating AI Clusters into Virtual Private Cloud

An architecture that decouples complex policy enforcement from high-speed packet forwarding to support VPC semantics on back-end NICs and enable front-end/back-end integration is proposed, suggesting that commodity hardware can support both high-throughput AI training and flexible VPC features.

Yinhe Wang, Xing Li, Enge Song et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.