Jul 2026· De Computis· Vol 15, pp. 439· 0 citations· 27 references
Computer Science
TL;DR
A language-independent benchmarking framework for evaluating fault tolerance in actor-based runtimes, which characterises supervised crash–recovery behaviour for largely stateless actor services rather than providing a comprehensive evaluation of actor-based fault tolerance.
Abstract
Fault tolerance is a fundamental requirement of distributed systems, and actor-based runtimes provide a widely adopted approach for building resilient and highly concurrent applications. Although several actor ecosystems offer mechanisms for supervision, failure detection, and recovery, comparative studies frequently focus on performance metrics rather than fault-tolerance behaviour. This paper presents a language-independent benchmarking framework for evaluating fault tolerance in actor-based runtimes. The framework was implemented using three representative ecosystems: Elixir/BEAM, Scala/Akka, and Go/Proto.Actor. A distributed chat-based benchmark application was used to measure throughput, reconnection latency, and failure-detection latency under recurring transient failures. All implementations followed an equivalent architecture and were executed under identical experimental conditions. The study deliberately targets a single, well-defined fault model: the supervised crash recovery of in-memory, effectively stateless actor services, in which chat actors are abruptly terminated and restarted by their supervisors while clients rediscover and reconnect to them. Stateful recovery (actor state, mailbox contents, in-flight or persistent messages), as well as multi-node network effects, are explicitly out of scope. Accordingly, the benchmark characterises supervised crash–recovery behaviour for largely stateless actor services rather than providing a comprehensive evaluation of actor-based fault tolerance. The results reveal distinct trade-offs among the evaluated ecosystems. Elixir achieved the highest throughput and the lowest throughput variability under fault conditions, while Scala/Akka consistently provided the lowest reconnection and failure-detection latencies, particularly at large scale. Go/Proto.Actor remained competitive in throughput-oriented scenarios but showed greater degradation in recovery-related metrics as concurrency increased. The results indicate that no single runtime dominates all evaluated dimensions of recovery behaviour. Beyond the runtime comparison, this work contributes a reproducible benchmarking framework that provides a foundation for future empirical studies of actor-based runtime recovery under controlled fault conditions.
Fault injection is widely used to evaluate the resilience of microservice systems, where client requests often span multiple services and execution stages. Existing request-level techniques usually control where and what faults are injected, but not when they are activated within a distributed execution. This limitatio...
Yuzhen Tan, Jian Wang, Bing Li et al.· arXiv.org· 0 citations
Distributed in-network programs are increasingly deployed in data centers for their performance benefits, but shifting application logic to switches also enlarges the failure domain. Ensuring their correctness before deployment is thus critical for reliability. While prior verification frameworks can efficiently verify...
Mingyuan Song, Huan-Xing Shen, Jinghui Jiang et al.· Conference on Applications,...· 0 citations
: Enterprise-scale distributed microservices in regulated environments operate under stringent availability, latency, and auditability requirements. Traditional monitoring approaches detect anomalies reactively, after service degradation has already impacted end users or regulatory SLA (Service Level Agreement) obligat...
Ketankumar Savajiyani· International Conference on...· 0 citations
Agent systems rely on LLM APIs for every response, but these APIs can return server errors, truncated responses, or corrupted content that propagates through downstream agents and causes task failure. Evaluating robustness under these faults is crucial for reliable deployment. Existing fault injection methods are offli...
Gou Tan, Zhensu Sun, Jieke Shi et al.· 0 citations
AGENTCHAOSBENCH is presented, a benchmark for detecting and localizing runtime faults in agentic systems from their execution telemetry, and its held-out labels and compact prediction format support reproducible comparison of LLM-based and non-LLM diagnosis methods.
Chenkai Zhang, Yiran Li, Yifang Tian et al.· 0 citations
DBA-Bench is presented, a benchmark addressing four gaps between evaluation and production operations: live-environment fidelity, outcome-first evaluation, and controlled scenario reproducibility, which uses instrumented PostgreSQL environments with active workloads, persistent state, and multi-source observations.