Skip to content
Open access

Transformer-Based Multi-Signal Predictive Autoscaling for SLA-Aware Resource Management in Kubernetes-Orchestrated Cloud-Native Environments

2024 · International Journal of Artificial Intelligence, Data Science and Machine Learning · Vol 5, pp. 241-249 · 0 citations

TL;DR

The study argues that SLA-aware predictive autoscaling should be treated not merely as a forecasting task but as an integrated control problem involving observability quality, model calibration, decision governance, and runtime safety.

Abstract

Kubernetes has become the dominant orchestration substrate for cloud-native applications, yet its native autoscaling mechanisms remain primarily reactive, threshold-driven, and limited in their ability to anticipate workload volatility before service-level agreement violations occur. Modern microservice systems exhibit non-linear interactions among request arrival rates, queueing delays, CPU saturation, memory pressure, network variability, pod cold-start latency, and downstream dependency bottlenecks. These characteristics make single-metric autoscaling policies insufficient for latency-sensitive workloads operating under strict service-level objectives. This paper proposes a Transformer-Based Multi-Signal Predictive Autoscaling framework for SLA-aware resource management in Kubernetes-orchestrated cloud-native environments. The proposed framework integrates heterogeneous observability signals, multi-horizon time-series forecasting, uncertainty-aware decision logic, and Kubernetes-native actuation to allocate resources before overload conditions materialize. Unlike conventional Horizontal Pod Autoscaler configurations that respond after resource utilization crosses predefined thresholds, the proposed approach forecasts near-future demand and performance risk using a Transformer encoder architecture designed to learn long-range dependencies, temporal seasonality, burst behavior, and cross-metric interactions. The framework translates predicted workload and latency risk into safe scaling actions through policy constraints that consider replica bounds, cooldown windows, pod readiness delays, cost budgets, and SLA violation probability. The paper develops the conceptual architecture, methodological workflow, evaluation metrics, and analytical discussion necessary for empirical implementation. The study argues that SLA-aware predictive autoscaling should be treated not merely as a forecasting task but as an integrated control problem involving observability quality, model calibration, decision governance, and runtime safety. The proposed model contributes to cloud resource management research by aligning deep temporal learning with Kubernetes operational semantics and by providing a structured pathway toward more reliable, efficient, and self-adaptive cloud-native platforms.

Read PDF

Similar papers

Conference Aug 2026

A Real-Time Workload Monitoring–Based Intelligent Auto-Scaling Framework for Cloud Systems

Kubernetes Horizontal Pod Autoscaler(HPA) and other existing auto-scaling solutions that respond reactively to demand experience significant delays in provisioning and inefficiencies when responding to sudden workload spikes. This paper proposes a new Real-Time Workload Monitoring-Based Intelligent Auto-Scaling Framewo...

Nikita Singh, Meenu Gupta, Rakesh Kumar et al. · 0 citations
Preprint Aug 2026

SLO-Scaler: Uncertainty-Aware SLO-Driven Autoscaling for Microservices

Autoscaling microservice-based applications to satisfy Service Level Objectives (SLOs) remains challenging due to bursty workloads, cascading latency across service dependencies, and cold-start overhead. Existing approaches such as the Kubernetes Horizontal Pod Autoscaler (HPA) rely on threshold-based CPU or memory met...

Shuo Wang, Xiao-Xuan Sun, Shao-yu Huang et al. · 0 citations
Open access Jul 2026

Event-Driven Autoscaling with KEDA and Karpenter: Cost Optimization and Elastic Throughput for Cloud-Native Workloads

Static threshold-based autoscaling with Kubernetes Horizontal Pod Autoscaler (HPA) does not cover asynchronous event-driven workloads that only target CPU and memory usage average metrics, such as the default HPA does. The HPA targets metrics that are lagging indicators‚ unlike the queue depth and stream backlog metric...

Avneet Bansal · 0 citations
Open access 2026

Counterfactual Autoscaling for Resource-Efficient Service Orchestration in the Cloud–Edge Continuum

Cloud–edge computing enables scalable and resilient deployment of microservice-based applications, however achieving resource efficiency while ensuring stringent Quality of Service (QoS) remains challenging. The strong interdependencies among microservices and non-linear latency effects near resource saturation render...

Lazaros Liatsas, Godfrey M. Kibalya, Angelos Antonopoulos · 0 citations
Open access Sep 2026

Cost-Efficient Predictive Auto-Scaling Using Transformer-LSTM Fusion Tuned with Bayesian Optimization

Cloud applications experience frequent and sometimes unpredictable shifts in demand due to user activity, daily usage cycles, and sudden workload spikes. Traditional autoscaling mechanisms used in cloud environments mostly follow reactive, threshold-based rules. They trigger scaling only after resources begin to satura...

K. V, A. V, Sivanantham S et al. · 0 citations
#edge computing Book Open access Sep 2026

QMScaler: A QoS-Constrained Resource-Efficient Microservice Autoscaling Framework for Edge Environments

QMScaler is proposed, a QoS-driven microservice horizontal scaling framework based on Monte Carlo Tree Search (MCTS) that aims to minimize the number of container instances while meeting the QoS requirements of multiple application functions.

Tianyang Zheng, Pengfei Yang, Zhe Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.