Skip to content
Open access

Cold Start Latency in Serverless Computing: Snapshotting, Fork-Based Models, and the Firecracker MicroVM Approach

Jul 2026 · Journal of Scientific Engineering Advances · Vol 2, pp. 1 · 0 citations · 1 references

TL;DR

The root causes are examined, a decomposition model is built, and three mitigation strategies are evaluated: memory snapshotting, fork-based execution (REAP), and Firecracker snapshot restore are evaluated.

Abstract

Cold start latency is a persistent, largely unsolved problem in serverless computing. When a function hasn't run recently, the platform must boot a VM or container, initialize a runtime, load the application framework, and only then handle the request — a process that takes 100ms to several seconds depending on platform and runtime. For latency-sensitive workloads, that overhead disqualifies the platform. We examine the root causes, build a decomposition model, and evaluate three mitigation strategies: memory snapshotting, fork-based execution (REAP), and Firecracker snapshot restore. We benchmark all three across five production-representative workloads. Snapshot restore cuts median cold start from 145ms to 28ms. Fork-based models (REAP) push that to 12ms. We also derive a warm pool sizing formula and model copy-on-write memory behavior under concurrent load. Our results show that fork-based approaches lead on latency while snapshot restore offers stronger isolation — making the right choice workload-dependent.

Read PDF

Similar papers

Open access Oct 2026

Arcus: Fast and Reliable Function State I/O for Serverless Computing With Log-Cache Co-Design

Serverless computing has emerged as a compelling cloud paradigm due to its simplified development model, automatic scalability, and fine-grained billing. While its stateless execution model enables high elasticity and resource efficiency, it poses noteworthy challenges for building complex stateful applications. To bri...

Yijie Liu, Zhuo Huang, Han-Xiang Huang et al. · 0 citations
Preprint Sep 2026

Reducing Cold-Start Latency in Serverless Applications via Dynamic Slicing

Across 31 applications on AWS Lambda, PyXtrim reduces cold-start latency by 21.7% and peak memory usage by 17.1% at the median, more than double the reduction achieved by the state of the art, while debloating each application in minutes.

Georgios Alexopoulos, Konstantinos Karakatsanis, Nikolaos Alexopoulos et al. · 0 citations
Preprint Aug 2026

Epico: Long-Lived WebAssembly Components for High-Performance Serverless Stream Processing

While serverless computing is popular, its dominant Function-as-a-Service (FaaS) model is ill-suited for stream processing because its stateless, centrally orchestrated functions cannot efficiently handle continuous, low-latency event flows. We introduce Epico, a serverless runtime explicitly designed to resolve these...

Matteo Della Bartola, Valerio Besozzi, Patrizio Dazzi et al. · 0 citations
Jul 2026

Fugue: Online Elasticity for Distributed Stateful Stream Processing

Fugue is a novel, self-contained reactive protocol that provides seamless and resource-efficient elasticity and reaches comparable handover performance while avoiding continuous replication overhead, and is implemented in Apache Flink.

Yu-Qiu Zhang, Yun-Hao Mao, Hans-Arno Jacobsen · 0 citations
Preprint Sep 2026

Gutenberg: Taming Latency-Critical Cloud Services with Near-Data-Processing

Gutenberg, a CPU+NDP for mutable, latency-critical cloud services, outperforms prior systems, reducing average and p99 latency by up to 80.4% and 85.8% and improves isolation and fairness while adapting to changing workload behaviors.

Qi Lin, Phillip B. Gibbons, Jovan Stojkovic et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.