Skip to content
Conference

Utility-Guided Orchestration for Cost-Efficient Tool-Augmented LLM Services

Jul 2026 · Fall Joint Computer Conference · pp. 333-338 · 0 citations · 31 references

Abstract

Tool-augmented large language model (LLM) services can solve complex tasks through retrieval and external tools, but current execution paradigms often trade adaptability for efficiency. Fixed workflows are predictable but rigid, while freeform reasoning loops such as ReAct may over-execute and issue redundant tool calls. We propose a lightweight utility-guided orchestration framework that formulates agent control as a costaware sequential decision problem over a compact action space: respond, retrieve, tool call, verify, and stop. An interpretable utility function balances expected gain, step-cost proxies, uncertainty, and redundancy. Experiments on multi-hop question answering show that the policy offers a controllable quality-cost trade-off and reduces token consumption by up to 10.6% in the semantic-redundancy setting while preserving similar answer quality. The framework is intended as an inspectable control layer for practical LLM services rather than a universally dominant accuracy optimizer.

View source

Similar papers

Conference Mar 2026

Utility-Guided Orchestration for Cost-Efficient Tool-Augmented LLM Services

A lightweight utility-guided orchestration framework that formulates agent control as a costaware sequential decision problem over a compact action space, intended as an inspectable control layer for practical LLM services rather than a universally dominant accuracy optimizer.

Bo-Yang Liu, Gongming Zhao, Hong-Liu Xu et al. · 4 citations
#machine learning Preprint Sep 2026

TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution

TRIAGE, a three-level routing framework that reduces token consumption by reusing historical execution trajectories, and proposes an automatic Skill extraction mechanism that distills high-frequency trajectory patterns into deterministic Skills, creating a positive feedback loop of the more you use it, the more efficie...

R. Wei · 1 citation
Jul 2026

Execution-First Synthetic Tool-Use Trace Generation for LLM Agents

SyntheticAgentTraceQA is proposed, an execution- first framework for generating scalable supervision data for tool- augmented agents and shows that execution-grounded supervision improves tool execution behavior, reference-trace agreement, and answer-generation performance on the evaluated tasks.

Hafsa Ouajdi, Francesco Giannuzzo, Alaa Boukhary et al. · 1 citation · ⚡1
#artificial intelligence Preprint Sep 2026

TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing

Agents tend to optimize, select, or constrain execution structures before decisive runtime outcomes are observed. However, such pre-execution commitment creates an orchestration bottleneck: when intermediate evidence invalidates the pending continuation, agents must either execute stale steps or replan broadly, compoun...

Tian-Xing Wang, Ming-Ming Zhao, Shuai Huang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses

LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed model: prompts, tool interfaces, middleware, state handling, and recovery logic. We study this setting as resource-bounded harness selection for fixed-model multi-turn tool agents, with the search surface scoped to prompt...

Cen-Mia Zhao, Hai-Bo Ruan, Wen-Jie Chen et al. · 3 citations · ⚡1
Jul 2026

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

E-Bench is introduced, a fully synthetic benchmark with 323 state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting, and it shows that multi-step tool use remains challenging: Pass^3 stays below 60% for the strongest models, and even with code execution in the E-Bench-Code exten...

Weihuang Zheng, Tianyuan Zou, Eileen Ye et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.