Skip to content
Review

AgentR A Stateful and Recovery-Aware Software Architecture for LLM-based Auditable Workflows

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

A stateful architecture for LLM workflow systems that enables improved observability, recoverability, and operational accountability in LLM workflow systems, and preliminary proof-of-concept that persistent state machine design, asynchronous orchestration, and cost-aware usage logging can enable improved observability, recoverability, and operational accountability in LLM workflow systems is provided.

Abstract

Modern LLM-based applications increasingly require multi- stage execution, persistent intermediate state, retry seman- tics, and auditable usage accounting. However, many LLM applications are still implemented as stateless prompt- response wrappers or session-bounded conversational sys- tems, which makes them difficult to recover, audit, and re- produce after interruption or failure. We propose AgentR, a stateful architecture for LLM workflow systems that en- ables persistence and recovery, instantiated through scien- tific literature review as a representative use case. AgentR represents research intent, generated queries, candidate- paper assessments and gap analyses as durable workflow artifacts, and executes the pipeline through asynchronous BullMQ workers backed by Redis, with PostgreSQL as the persistence store. The design includes explicit processing state transitions, retries with exponential backoff, orphan job detection, credit-aware pre-checks, ACID token-cost logging, and Type-2 slowly changing pricing records. We evaluate AgentR on telemetry collected from a prototype deployment. At the LLM stage, the system achieves 99.2% job completion, and mean latencies of 9.0 s, 18.9 s, and 25.4 s for intent decomposition, query generation, and paper scoring, respectively. Parallel scoring allows for analytical latency modeling from observed calls, leading to as much as 4.3 wall-clock speedup over sequential execution. The results provide preliminary proof-of-concept that persistent state machine design, asynchronous orchestration, and cost-aware usage logging can enable improved observability, recoverability, and operational accountability in LLM workflow systems. The prototype implementation of AgentR is publicly available at: https://github.com/ RiyaSamanta/AgentR-public.

View source

Similar papers

Open access Sep 2026

ReliHarness: A Self-Learning Reliability Harness for LLM Agent Tool Execution

Large language model agents invoke external tools through the Model Context Protocol to perform actions beyond text generation. Once a background thread is dispatched, even the latest Model Context Protocol Tasks extension exposes only coarse-grained, self-reported task states, so a thread that deadlocks or crashes sil...

Wen-Zhi Chen, Peng-Han Song, Qi-Chao Lu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

Distributed LLM-agent teams can read the latest shared facts and still act on an obsolete plan. A planner may derive an action from requirement $r_3$, another agent may commit $r_4$, and an executor may receive $r_4$ without replacing the plan derived from $r_3$. We call this \emph{stale-plan execution}: state freshnes...

Evan Chen, Shi-Qiang Wang, Christopher G. Brinton · 0 citations
Open access Aug 2026

From Logging Configuration to Code Execution: A Systematization of Log4j 2 File-Write Primitives in HTTP-Exposed JMX

Java middleware may expose Java Management Extensions (JMX) through Jolokia’s Hypertext Transfer Protocol (HTTP) bridge. In affected ActiveMQ deployments, reachable Log4j 2 configuration managed beans (MBeans) become write capabilities and, with compatible triggers, enable remote code execution (RCE). We ask: in a spec...

A. Caciulescu, Matei Badanoiu, R. Rughinis et al. · 0 citations
Preprint Aug 2026

When"Must"Becomes"Maybe": Constraint Weakening in LLM Agent Workflows

This work identifies a state-transmission failure between information extraction and action in large language model agents, and shows how handoff transformations can retain state content while weakening its constraints on downstream action.

Yi-Heng Sun, Huifei Wang, Yan-Cheng Zhu et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.