Skip to content

DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving

Sep 2026 · 0 citations · 112 references
Computer Science

TL;DR

DynBranch is proposed, which makes an unresolved branch addressable before it resolves, and its stable coordinate lets candidate subgraphs run during resolution and completed subgraph results be reused across later requests.

Abstract

Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the branch-resolution barrier. Caching alone does not hide it: the key that identifies a reusable result is not known until then. In this paper, we propose DynBranch, which makes an unresolved branch addressable before it resolves. Its stable coordinate lets candidate subgraphs run during resolution and completed subgraph results be reused across later requests. A two-level controller admits this work when its expected benefit exceeds the load price. DynBranch sits at the model-API boundary and requires no changes to agent harnesses or model execution engines. Across four agentic workloads with Qwen3-32B on 4x H200 GPUs, DynBranch reduces mean latency by up to 32% over each workload's strongest prior system and by 46-66% against a no-reuse floor, while preserving workflow results. The benefit persists across backbone families and on a commodity Qwen3-8B/RTX 4090 deployment.

View source

Similar papers

Preprint Sep 2026

PipeSwift: Revisiting Pipeline Parallelism for Large-Scale Completion-Oriented Agentic LLM Serving

It is shown that pipeline parallelism (PP), long overlooked because it offers little decode-latency advantage, can reduce JCT by providing a more favorable balance between prefill and decode efficiency, and PipeSwift is built, an optimized open-source pipeline-parallel runtime integrated with a tailored micro-batch par...

Shiju Wang, Fei Ren, Fang-Cheng Fu et al. · 0 citations
#natural language process... Preprint Sep 2026

TomasuLLM: Out-of-Order Speculative Execution for LLM Agents

Long-running tools can dominate coding-agent latency: compilers, test suites, and repository commands take seconds to minutes while the agent idles. This observation stall presents the same tension that drove out-of-order processors -- asequential interface hides work that can be predicted and started early, but a spec...

Jiang-Nan Yu, Ce-Yu Xu, Meng-Ming Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LazyAgent: Demand-Driven Materialization and Physical Optimization of Agentic Programs

Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present LazyAgent, a unified execution framework for agent-authored programs organized around a live, goal-derived demanded set. LazyAgent refreshes a backward closure from requested outputs as execution state changes and...

Xin Heng · 0 citations
Preprint Sep 2026

AKTS: Sub-Microsecond Kernel Policy Switching for Language-Model Agents

GPU-backed LLM servers often multiplex interactive requests with background batch work on the same CPUs. During a request burst, the scheduler should protect time-to-first-token; between bursts, it should let background work make progress. A fixed kernel policy leaves one of these objectives on the table, so agentic OS...

Mohammadali Khodabandehlou, Mahdi Alizadeh · 0 citations
Preprint Aug 2026

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

PeakBench is a benchmark of executable multi-tool workflows with execution-grounded dependency annotations and measured resource profiles that shows that strong logical planning does not reliably translate into safe or efficient execution under resource constraints, and exposes resource information to reduce avoidable...

Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li et al. · 0 citations
#machine learning Preprint Sep 2026

EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the...

Kun-Ming Shao, Jie-Run Chen, Jiang-Nan Yu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.