Skip to content

Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

The latency with respect to the resource dynamics of processing multiple requests and tasks concurrently is characterized and new optimization opportunities that exploit the resource dynamics of tasks are demonstrated: CPU-aware tool admission and task-aware CPU allocation.

Abstract

LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers. This execution model can make the optimization of agent serving difficult because latency, local resource demand, and container bottlenecks inter-mix across requests. However, the current agent ecosystem runs without much consideration of resource dynamics, which results in significant waste of the precious resources. This paper analyzes the resource inter-mix of AI agents for three representative tasks: retrieval-augmented question answering, web search, and software coding. To this end, we characterize the latency with respect to the resource dynamics of processing multiple requests and tasks concurrently. Our measurements show that agents have a wide range of behaviors depending on tasks, so that even the same tool can differ substantially in resource dynamics. We also find that running multiple requests concurrently exposes task-dependent bottlenecks in resource dynamics such as CPU, disk I/O, and memory. Furthermore, we uncover that faster LLM responses or more CPU cores do not always accelerate agents. Based on these observations, we demonstrate new optimization opportunities that exploit the resource dynamics of tasks: CPU-aware tool admission and task-aware CPU allocation. Our results show that the latency of CPU-sensitive agent tasks improves $\sim$5.4$\times$, and the average latency across multiple tasks is reduced $\sim$32% compared to native agents.

View source

Similar papers

Preprint Sep 2026

AgentReplay: Token-Wise Trace Replay Is Essential for Fair Serving System Performance Benchmarking

LLM-based agents execute multi-turn workflows with interleaved model inference and tool calls, making efficient serving increasingly important. However, evaluating serving optimizations is challenging because identical tasks can produce different execution trajectories. Changes in generated tokens can alter subsequent...

Zai-Feng Pan, Michael Wang, Chris Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Substrate-Aware AI Agents: Execution Context as a First-Class Input

Autonomous AI agents increasingly select actions in environments whose memory, execution-time, runtime, compute, and operational constraints determine what counts as a suitable plan. We call the absence of this execution context from an agent's planning state substrate blindness. We test this general proposition throug...

Manu Agrawal · 0 citations
Preprint Aug 2026

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

This work presents AgentSysBench, a benchmark suite and measurement toolkit with ten representative agentic applications and unified systems-level instrumentation, and identifies six properties that distinguish agentic workloads from conventional LLM serving.

Chaokun Chang, Yu-Kun Zhou, Kai-Hua Fu et al. · 8 citations
Open access Aug 2026

An Explainable Agentic AI Framework for Intelligent Multi-Cloud Resource Allocation

An explainable agentic AI framework for multi-cloud task allocation built on a contextual-bandit agent that observes each provider's current price, estimated latency, and load before autonomously selecting a placement, then updates its policy online from the resulting cost, latency, and service-level-agreement (SLA) ou...

Dr. Sajitha A V · 0 citations
#artificial intelligence Preprint Sep 2026

Token Efficient Task Execution via Application Behavior Modeling for Web Agents

The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web-application tasks described in natural language, by analyzing the web-applicatio...

Alexandru Ianta, Eleni Stroulia · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.