Skip to content

SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision

May 2026 · arXiv.org · Vol abs/2606.01139 · 17 citations · 52 references
Computer Science

TL;DR

Evaluated across three main benchmarks, two domain-specific studies, and six LLMs, SkillRevise substantially outperforms one-shot baselines, and the revised skills transfer across both executors and task environments, suggesting that SkillRevise captures reusable procedural knowledge beyond any single executor.

Abstract

Agent skills are procedural artifacts that enable LLM agents to execute workflows, verify constraints, and recover from failures. Existing self-evolving methods refine skills using accumulated trajectories. However, they struggle in cold-start settings, where only an initial, imperfect skill is available. Consequently, skill construction defaults to expert authoring or one-shot LLM generation. Expert-authored skills are costly and may not align with how LLM agents actually execute tasks, while one-shot generated skills can be syntactically well formed yet behaviorally weak. To bridge this gap, we propose SkillRevise, an execution-grounded framework designed to iteratively refine these initial skills. SkillRevise diagnoses skill defects from execution evidence, retrieves relevant repair principles from a general memory, and applies execution-anchored edits. By re-executing candidates and measuring empirical utility, it retains the best observed skill within the revision budget. Evaluated across three main benchmarks, two domain-specific studies, and six LLMs, SkillRevise substantially outperforms one-shot baselines, improving the base agent's success rate on SkillsBench from 36.05% to 61.63%. Furthermore, the revised skills transfer across both executors and task environments, suggesting that SkillRevise captures reusable procedural knowledge beyond any single executor. Our code is available at https://github.com/HKUST-KnowComp/skillrevise.

View source

Similar papers

Preprint Aug 2026

SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance

SkillSentry is proposed, a skill-oriented runtime assurance framework built upon a new domain-specific language (DSL) for representing runtime guidance for skill execution that improves the task success rate of LLM agents by 24.1% across skills, on average, while exhibiting lower variability across repeated runs.

You Lu, Xinyu Huang, Bi-Huan Chen et al. · 1 citation
Preprint Aug 2026

Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following procedural guidance and improving it from execution evidence are distinct capabilities. Inference time loops can repair skills but do not improv...

Kang Peng, Zhi-Wei Zhang, Yichen Zhang et al. · 1 citation
#artificial intelligence Review Sep 2026

SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose Ski...

Renxi Wang, M. Hee, Fajri Koto et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

Textual skills enable large language model (LLM) based agents to accumulate reusable procedural knowledge without updating model parameters. Yet existing skill evolution remains largely confined to the text space: an optimizer must diagnose success and failure patterns, and revise skills solely from long execution traj...

Kai-Xin Zhang, Chang-Ming Li, Ying-Dong Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SkillAlign: Aligning Skill Interfaces for LLM-based Agents

SkillAlign is proposed, a provider-agnostic framework that represents candidate skills as multi-view procedural cards and renders them through alternative exposure interfaces, including full instructions, hints, compressed summaries, workflows, or no exposure, which enables counterfactual evaluation where the task, age...

Shuo Ren, Xiaomian Kang, Jia-Jun Zhang · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.