Skip to content

Field-Aware Agent Skill Retrieval

Aug 2026 · 0 citations · 11 references
Computer Science

TL;DR

The results show that skill representation itself matters, and that simply preserving the structure already present in skill files can substantially improve retrieval.

Abstract

As lifelong learning agents accumulate lifelong growing skill banks, retrieving the correct skill becomes an increasingly important bottleneck. Most current skill retrieval methods treat each skill as one flat document by concatenating fields such as the name, description, and body. However, skills are naturally structured, multi-field objects, where each field provides different information about when and how the skill should be used. In this work, we study whether preserving this structure improves skill retrieval. We represent each skill as its separate components, and compute sparse and dense similarities for each field independently, exposing a naturally tensorized, field-aware representation of the skill bank. We then combine these field-level scores either with uniform weights or with a small learned MLP. Across two different skill retrieval benchmarks, SkillRet and SRA-Bench, we find that keeping fields separate improves hybrid retrieval, and learning over the field-level scores gives the strongest and most consistent results. Our field-aware MLP reaches $77.95$ Recall@10 on SkillRet and $83.78$ Recall@10 on SRA-Bench, outperforming the corresponding concatenated learned baselines. We also find that the advantage grows as the skill bank becomes larger, suggesting that field-aware skill retrieval becomes especially useful in the setting where retrieval is most difficult. Our results show that skill representation itself matters, and that simply preserving the structure already present in skill files can substantially improve retrieval.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated re...

Guanqun Yang, Wen-Long Zhang, Tian Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SkillFM: Generating Skills for LLM Agents via Latent Flow Matching

Textual skills provide reusable guidance for large language model agents, but existing approaches often rely on manually curated skill banks or reinforcement learning with indirect and delayed feedback. We introduce SkillFM (Skill Flow Matching), a generative framework that synthesizes task-conditioned textual skills d...

Zu-Ming Zhang, Jie He, Yi-Zhe Zhang et al. · 0 citations
#natural language process... Preprint Sep 2026

M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use

Agent skills, reusable procedural documents that extend LLM agents beyond their parametric memory, have become an important interface for deploying agents on real-world tasks. Community-maintained skill libraries built around this interface are growing rapidly. However, this ecosystem remains deeply English-centric: ou...

Yi-Lun Liu, Shi-Min Tao, Ming-Gui He et al. · 0 citations
Preprint Aug 2026

Skills Know Their Neighbors: Cluster-Contrastive Capability Pages for Skill Retrieval

As skill libraries grow, large language model agents must retrieve reusable skills from candidates that often share the same topic and vocabulary but implement different capabilities. Retrieval is limited not only by the scorer but also by the text being scored: a document may describe what a skill does without stating...

Zifei Wang, Wei Wen, Qian Ji et al. · 4 citations · ⚡1
Preprint Aug 2026

SkillEval: Decomposing Agent Skill Quality into Interpretable Signals

SkillEval is used to evaluate skills in controlled quality tests and it is used for diagnosing weaknesses in skill documents and guiding targeted revisions, and it is shown that SkillEval reliably distinguishes skills of different quality.

Jia-Hui Han, Qinuo Li, Ziheng Peng et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents

Large language model (LLM) agents increasingly rely on external skills, but routing user requests over large skill registries is difficult because many skills are functionally redundant while complex tasks often require complementary skill sets. Existing skill routers typically rank candidates independently by query re...

Wang Wei, Tiankai Yang, Samyadeep Basu et al. · 2 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.