Skip to content

ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

Sep 2026 · 0 citations · 54 references
Computer Science

TL;DR

This work proposes ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection, which introduces category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools.

Abstract

Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively search, distinguish, and compose tools under context-length constraints. We identify large-scale tool selection as a new challenge for agentic reinforcement learning, highlighting that existing RL methods for knowledge-based question answering are inadequate for selecting tools while considering compatibility. To address this challenge, we propose ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection. Specifically, we introduce category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools, event-level search modeling to explicitly optimize the discovery of target tools during multi-turn search, and trajectory-aligned credit allocation to provide fine-grained reward signals for different stages of the search-selection process. Extensive experiments on large-scale tool selection benchmarks demonstrate that ToolSearcher consistently outperforms a set of strong baselines in challenging settings involving iterative search and complex tool composition.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Toolcompass: Guiding Tool Trialing, Not Suppressing It

This work introduces ToolCompass, a post-training framework that guides tool trialing by organizing tool-call representations according to shared functions and jointly reduces intra-function variation across domains and increases inter-function separation.

Jun-Lin Fang, Chong-Chong Zhang, Do Nguyen-Thanh et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Self-Evolving Agents via Likelihood-Guided Tool-Space Optimization

Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisions. We study tool-space self-evolution,...

Xuanqi Zhang, Rui-Nan Jin, Run Yang et al. · 0 citations
Preprint Aug 2026

Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents

Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external observations rather than model parameters alone. This creates an evaluation mismatch for LLM unlearning: previous unlearning methods may suppress direct parametric recall, but an agent...

Baicheng Chen, Zhe-Yuan Liu, Jingyu Zhang et al. · 0 citations
#machine learning Preprint Aug 2026

One Policy Is Enough: Single-Agent Reinforcement Learning Outperforms Tree Search for Chemistry Tool Learning

This model interleaves reasoning, tool calls, and returns in one left-to-right generation, trained by a supervised warm-up and then outcome-level reinforcement learning against a programmatic reward read directly off the gold call chain, which leaves no learned critic and no judge in the training loop.

Armin Dariani, Sifan Wu, Bang Liu et al. · 0 citations
Preprint Aug 2026

Joint Optimization of Tool Creation and Use for Large Language Model Agents

A reinforcement learning framework that jointly trains tool creation and tool use inside a single policy, with three separate reward axes that catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient.

Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen et al. · 2 citations
Preprint Aug 2026

MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning

Reinforcement learning (RL) has become an effective way to improve the tool-use ability of large language models (LLMs), but most existing RL frameworks stop at the policy update. For every new domain, the user is left with two hard systems problems: standing up an isolated environment for each of hundreds of concurren...

Ziyang Luo, Yan Yang, Xiang-Ru Jian et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.