Skip to content
Conference

Challenging the Two-Core Assumption: Deterministic Single-Core Zephyr Virtualization with PCIe NIC Passthrough

Aug 2026 · IEEE International Conference on Embedded and Real-Time Computing Systems and Applications · pp. 201-208 · 0 citations · 22 references

Abstract

Industrial edge platforms increasingly consolidate real-time control and general-purpose workloads on a single system-on-chip (SoC) to reduce costs, power, and complexity. Conventional real-time virtual machine (RTVM) setups, however, typically rely on PREEMPT_RT Linux and commonly reserve at least two CPU cores to isolate real-time tasks from housekeeping, which hinders scalability on resource-constrained edge platforms. This paper investigates whether low-latency Ethernet networking with tight observed tail latency can be sustained in a single-core RTVM configuration. We implement a single-core Zephyr-based RTVM on ACRN and compare it with a single-core PREEMPT_RT Linux RTVM under the same VM topology, using the same directly assigned Intel i226-LM PCIe Ethernet NIC via passthrough and an identical UDP echo workload. Latency measurements across 30 million packets at a traffic rate of 8 thousand packets per second (8 kpps) characterize both averagecase and extreme-tail behavior. PREEMPT_RT Linux shows severe tail amplification even without interference (99.999th percentile: $\mathbf{2 1 6 3} \boldsymbol{\mu} \mathbf{s}$; max: $\mathbf{6 8 7 2} \boldsymbol{\mu} \mathbf{s})$, while Zephyr maintains a tightly concentrated latency distribution (99.999th percentile: $79 \mu \mathrm{s}$; max: $82 \mu \mathrm{s})$. Under full-system noisy-neighbor load, Zephyr preserves sub- $\mathbf{1 0 0}-\boldsymbol{\mu} \mathbf{s}$ observed tail latency, whereas Linux degrades further. These empirical findings provide evidence that a specialized RTOS-based RTVM can sustain tight tail-latency performance and low delay variation on a single core, challenging the conventional two-core provisioning strategy for real-time edge systems.

View source

Similar papers

Preprint Aug 2026

Offering Microsecond-Scale Cross-VM Core Elasticity on Colocated Lightweight Virtual Machines

Serverless platforms commonly colocate many diverse workloads, each in a fast-booting, memory-lean virtual machine (VM), to improve deployment density. Overprovisioning each VM for its peak protects tail latency during traffic bursts but hurts density; maintaining high density while effectively protecting tail latency...

Yi-Bo Yan, Seo Jin Park · 1 citation
Conference Aug 2026

Characterizing Predictability–Latency Trade-offs of KV-Cache SSD Offloading in LMCache for LLM Serving Systems

KV-cache offload is widely used to stretch GPU memory for LLM serving, but its storage behavior has not been characterized at the block-device level. In this paper, we study LMCache through realworld multi-session workloads that span same/different context $\times$ same/different prompt, using over 100 stateless reques...

Ying He, Dingsen Shi, Yanbo Dai et al. · 0 citations
Open access Aug 2026

Design of a Multi-Tenant Real-Time Inference Framework Based on OpenStack and SR-IOV GPU Virtualization

A standards-based, multi-tenant cloud inference framework that integrates OpenStack orchestration with Single Root I/O Virtualization (SR-IOV)-enabled graphics processing unit (GPU) partitioning to achieve predictable and isolated real-time inference execution.

Rui Ma, Bing-Feng Shi, Xing-Run Ma et al. · 0 citations
Preprint Aug 2026

PANEM: A Heuristic Latency Model

PANEM is presented, a lightweight event-driven heuristic model that has guided four generations of commercial server-core development at Ampere Computing and shows that a calibrated, contention-aware abstraction can deliver practical predictive value for industrial design-space exploration at simulation costs similar t...

Lucas Crowthers, M. Madhav · 1 citation
Open access Sep 2026

A Deterministic FPGA-CPU Data Interaction Architecture with Cache-Aware Core Selection for Real-Time Systems

Achieving periodic, low-latency data transfer with deterministic timing guarantees between I/O devices (e.g., Field-Programmable Gate Arrays (FPGAs)) and hosts remains a fundamental challenge in real-time domains, including automated driving, high-frequency trading, and hardware-in-the-loop systems. Unpredictable commu...

Hui-Ji Zheng, Qing Liu, Kang Wang et al. · 0 citations
Preprint Aug 2026

PRO-RAN: Processor-Level Characterization of Open RAN Centralized and Distributed Units

Open Radio Access Network (O-RAN) disaggregates RAN protocol functions and enables Centralized Unit (CU) and Distributed Unit (DU) software to execute on general-purpose computing platforms. Different CU and DU protocol responsibilities produce different processor workloads and execution paths. Conventional performance...

Moojan Kamalzadeh, Larry J. Horner, Linqi Xiao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.