Skip to content
Book Open access

Krysha: Cost-Efficient Resource Orchestration for Geo-Distributed Serverless Microservices

Jul 2026 · IEEE International Symposium on High-Performance Parallel Distributed Computing · pp. 319-333 · 0 citations · 93 references
Computer Science

TL;DR

Krysha is presented, an adaptive orchestration framework that jointly optimizes function scheduling and resource allocation for geo-distributed serverless microservices and can achieve up to 74.7% cost savings in scaled deployments compared to state-of-the-art alternatives while maintaining SLO requirements.

Abstract

The convergence of microservice architectures and serverless computing promises an elastic and cost-efficient model for modern cloud applications that often span multiple geo-distributed regions. However, prevailing serverless orchestrators that prioritize resource utilization or simple cold-start mitigation often prove suboptimal concerning SLO compliance and cost-efficiency in this emerging use case. In this paper, we present Krysha, an adaptive orchestration framework that jointly optimizes function scheduling and resource allocation for geo-distributed serverless microservices. Krysha employs a novel bi-level scheduling strategy: global-level early-binding to regions for fast function dispersion, coupled with regional-level late-binding to compute nodes for optimized resource use and cost. Moreover, Krysha achieves fine-grained resource allocation by decoupling CPU and memory provisioning and applying in-place vertical scaling on individual function instances. These capabilities are guided by a comprehensive cost model and practical online optimization techniques. Our extensive evaluation shows that Krysha can achieve up to 74.7% cost savings in scaled deployments compared to state-of-the-art alternatives while maintaining SLO requirements.

Read PDF

Similar papers

2026

Resource Allocation and Container Scaling for Microservices in Multi-Cluster Edge Computing System

With the advent of the 6G era and the evolution of distributed systems, edge computing has become a pivotal architecture for deploying latency-sensitive, resource-efficient applications. In particular, the microservice architecture, characterized by modular and loosely coupled components, has gained significant tractio...

Jing-Yang Voon, Yao Chiang, Hung-Yu Wei · 0 citations
#reinforcement learning Open access Sep 2026

CoSIPR: Shared Service Orchestration with Dynamic Interest Coalitions in Edge Computing

Resource-intensive mobile edge computing (MEC) services are often provisioned on a per-request basis, resulting in repeated activation of equivalent service instances and redundant transmission of the same category-level state over overlapping inter-station links. Existing approaches rarely integrate demand aggregation...

Meng-Xuan Dai, Xuan Chen, Ling Yang et al. · 0 citations
Jul 2026

Multi-objective formulation of efficient microservice deployment in Kubernetes with dynamic resource allocation using hybrid nature-inspired algorithm

The primary innovation of this research is the development of an adaptive switching framework that integrates Clouded Leopard Optimization for robust global exploration with Cock-hen-chicken Optimization for hierarchical local refinement.

Srinivasan Lingaraj, Purushothaman Annadurai · 0 citations

CERLA-SFC: hierarchical orchestration of cost-efficient, reliable and low-latency SFCs in multi-access edge computing

CERLA-SFC is introduced, a hierarchical, multi-objective orchestrator that unifies learning-based placement, topology-aware routing and event-driven resource allocation in a single control loop that maintains near-zero latency violations across all urgency classes while keeping end-to-end delay in the millisecond range...

Yuanfei Xiao, Zhenli He, Xiaolong Zhai et al. · 0 citations
Open access 2026

SLA-DE-RALBA: Cost-efficient dynamic enhanced resource-aware load balancing algorithm for cloud computing

An SLA-aware Dynamic Enhanced Resource-Aware Load Balancing Algorithm (SLADE- RALBA) that minimizes load imbalance by considering the computational capacities of virtual machines and ensures Service Level Agreement (SLA) compliance through a three-tier priority-based workflow is proposed.

Mohsin Nawaz, Altaf Hussain, Marran Al Qwaid et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.