Skip to content

SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

Sep 2026 · 0 citations · 38 references
Computer Science

TL;DR

This work argues the missing unit of reuse is the solving procedure shared by a cluster of related tasks, and builds SkillGLoW (Global-Local Weave) around it: the local skills a task writes from its own execution are aggregated into procedural families and compressed into de-instantiated global priors.

Abstract

LLM agents increasingly self-improve by writing and reusing textual skills, kept either as one global document or as a flat pool of per-task entries, though most of the evidence comes from domains with structurally similar tasks. On long-horizon workloads where each task demands a different solution, the two forms fail in opposite ways: the document collapses into generic discipline, while the pool inflates and its entries stay bound to the instance that wrote them. We argue the missing unit of reuse is the solving procedure shared by a cluster of related tasks, and build SkillGLoW (Global-Local Weave) around it: the local skills a task writes from its own execution are aggregated into procedural families and compressed into de-instantiated global priors, while the instance detail they hold is regenerated per task rather than stored; a commit gate admits a prior only when real execution shows it does not degrade the deployed library. Across four benchmarks (mathematical reasoning, terminal automation, software repair, and embodied control) and three models, the priors gain 17.2 points (hard) over the no-skill baseline on average, with positive gains in all 12 continual-improvement runs, and 18.0 with local regeneration, while the library holds one prior per procedural family, 3.6x more compact than the per-task pool. Under the same protocol GLoW leads a published single-document optimizer on 15 of 21 cells. Unmodified, the library lifts success on unseen ALFWorld tasks from 73.9% to 83.9%, evidence that what transfers is procedure rather than task memory.

View source

Similar papers

Preprint Aug 2026

SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent

SkillProx is introduced, a proximal-gradient-inspired forward--backward framework that couples closed-loop diagnostic evolution with utility-aware proximal refinement and demonstrates the complementary effects of closed-loop diagnosis and proximal refinement.

Mingxuan Zheng, Yu-Jin Zhou, Chuxue Cao et al. · 2 citations
Preprint Aug 2026

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

SkillZip is presented, an evaluation-free method that compresses a skill by finding its shortest faithful structural explanation, subject to a hard coverage constraint for every extracted trigger, workflow edge, tool requirement, obligation, and output field.

Xiao-Fan Bai, Hong-Qiang Lin, Chao Liu et al. · 2 citations
#artificial intelligence Preprint Aug 2026

FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

Agents used over time encounter recurring professional work: each case requires different evidence and judgment, while the underlying workflow can be reused. Benchmarks built from independent tasks cannot reveal whether an agent turns earlier experience into better procedures for later cases. We introduce FinEvo-Bench,...

Bo Deng, Kang Zhou, Li-Fan Guo et al. · 1 citation
#artificial intelligence Preprint Sep 2026

SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated re...

Guanqun Yang, Wen-Long Zhang, Tian Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SkillFocus: Evolving Agent Skills via Capability Decomposition

Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision to the behavior of th...

Ning Wang, Z. Gong, Bing-Dong Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SkillAlign: Aligning Skill Interfaces for LLM-based Agents

Language-model agents increasingly rely on skills: reusable procedural knowledge for reasoning, tool use, and interaction. Existing work studies how skills are acquired, retrieved, compressed, or composed, but often assumes that once a skill is selected, its interface to the agent is fixed. We argue that this overlooks...

Shuo Ren, Xiaomian Kang, Jia-Jun Zhang · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.