Skip to content
Open access

Benchmarking Fault-Tolerance Characteristics of Actor-Based Runtimes

Jul 2026 · De Computis · Vol 15, pp. 439 · 0 citations · 27 references
Computer Science

TL;DR

A language-independent benchmarking framework for evaluating fault tolerance in actor-based runtimes, which characterises supervised crash–recovery behaviour for largely stateless actor services rather than providing a comprehensive evaluation of actor-based fault tolerance.

Abstract

Fault tolerance is a fundamental requirement of distributed systems, and actor-based runtimes provide a widely adopted approach for building resilient and highly concurrent applications. Although several actor ecosystems offer mechanisms for supervision, failure detection, and recovery, comparative studies frequently focus on performance metrics rather than fault-tolerance behaviour. This paper presents a language-independent benchmarking framework for evaluating fault tolerance in actor-based runtimes. The framework was implemented using three representative ecosystems: Elixir/BEAM, Scala/Akka, and Go/Proto.Actor. A distributed chat-based benchmark application was used to measure throughput, reconnection latency, and failure-detection latency under recurring transient failures. All implementations followed an equivalent architecture and were executed under identical experimental conditions. The study deliberately targets a single, well-defined fault model: the supervised crash recovery of in-memory, effectively stateless actor services, in which chat actors are abruptly terminated and restarted by their supervisors while clients rediscover and reconnect to them. Stateful recovery (actor state, mailbox contents, in-flight or persistent messages), as well as multi-node network effects, are explicitly out of scope. Accordingly, the benchmark characterises supervised crash–recovery behaviour for largely stateless actor services rather than providing a comprehensive evaluation of actor-based fault tolerance. The results reveal distinct trade-offs among the evaluated ecosystems. Elixir achieved the highest throughput and the lowest throughput variability under fault conditions, while Scala/Akka consistently provided the lowest reconnection and failure-detection latencies, particularly at large scale. Go/Proto.Actor remained competitive in throughput-oriented scenarios but showed greater degradation in recovery-related metrics as concurrency increased. The results indicate that no single runtime dominates all evaluated dimensions of recovery behaviour. Beyond the runtime comparison, this work contributes a reproducible benchmarking framework that provides a foundation for future empirical studies of actor-based runtime recovery under controlled fault conditions.

Read PDF

Similar papers

Jul 2026

SequenceFI: Non-intrusive Temporal Fault Injection for Microservice Systems

Fault injection is widely used to evaluate the resilience of microservice systems, where client requests often span multiple services and execution stages. Existing request-level techniques usually control where and what faults are injected, but not when they are activated within a distributed execution. This limitatio...

Yuzhen Tan, Jian Wang, Bing Li et al. · 0 citations
Book Open access Aug 2026

Towards Efficient Verification of Distributed In-Network Computing Programs

Distributed in-network programs are increasingly deployed in data centers for their performance benefits, but shifting application logic to switches also enlarges the failure domain. Ensuring their correctness before deployment is thus critical for reliability. While prior verification frameworks can efficiently verify...

Mingyuan Song, Huan-Xing Shen, Jinghui Jiang et al. · 0 citations
Open access 2026

Agentic AI for Self-Healing Microservices: An LLM-Orchestrated Framework for Autonomous Fault Detection and Remediation in Regulated Enterprise Environments

: Enterprise-scale distributed microservices in regulated environments operate under stringent availability, latency, and auditability requirements. Traditional monitoring approaches detect anomalies reactively, after service degradation has already impacted end users or regulatory SLA (Service Level Agreement) obligat...

Ketankumar Savajiyani · 0 citations
Preprint Aug 2026

AgentChaos: Chaos Engineering for Agent Systems via Programmatic Fault Injection

Agent systems rely on LLM APIs for every response, but these APIs can return server errors, truncated responses, or corrupted content that propagates through downstream agents and causes task failure. Evaluating robustness under these faults is crucial for reliable deployment. Existing fault injection methods are offli...

Gou Tan, Zhensu Sun, Jieke Shi et al. · 0 citations
Preprint Aug 2026

When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry

AGENTCHAOSBENCH is presented, a benchmark for detecting and localizing runtime faults in agentic systems from their execution telemetry, and its held-out labels and compact prediction format support reproducible comparison of LLM-based and non-LLM diagnosis methods.

Chenkai Zhang, Yiran Li, Yifang Tian et al. · 0 citations
Jul 2026

DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents

DBA-Bench is presented, a benchmark addressing four gaps between evaluation and production operations: live-environment fidelity, outcome-first evaluation, and controlled scenario reproducibility, which uses instrumented PostgreSQL environments with active workloads, persistent state, and multi-source observations.

Jun-Ming Chen, Jun-Yang Jiang, Xu Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.