Skip to content

Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents

Sep 2026 · 0 citations
Computer Science

TL;DR

This work mine harder tasks --- merged pull requests of the repository, reverted at a single frozen base commit; and score a candidate document by whether the same agent does better with it than without it, to find a maintainer of one repository found in them knowledge one only gets by working in the project.

Abstract

Coding agents increasingly read repository knowledge from SKILLs --- plain \texttt{.md} files versioned alongside the code. Recent work synthesizes these files automatically, by optimizing the document against a benchmark. A bare repository comes with no benchmark, and the synthetic tasks prior work builds are small enough that a capable agent saturates them with no document at all. We mine harder tasks --- merged pull requests of the repository, reverted at a single frozen base commit; and score a candidate document by whether the same agent does better with it than without it. On three Kotlin repositories, the documents GEPA finds raise this score by $4.9$pp on average, and the ones SkillOpt finds leave it where it started, $0.1$pp above the seed. The GEPA gain matches what prior work reports with the same optimizer, and at the dataset size a single repository supplies it cannot be separated from the agent's run-to-run variance; settling that would take more tasks than one repository's history yields. The documents themselves read better than the score: a maintainer of one repository found in them knowledge one only gets by working in the project.

View source

Similar papers

Book Open access Sep 2026

Context Inflation: The Hidden Cost of AI Coding Agents in Real-World Repositories

A growing narrative holds that AI coding agents can now do the work of junior software developers, weakening the case for hiring and training them. The evidence offered is almost entirely benchmark performance: leaderboards like SWE-bench report not only how often an agent resolves a task but the dollar cost of each re...

Tommaso Turchi, Lorenzo Palazzo · 0 citations
Preprint Aug 2026

The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks

This work reconstructs a coupled-fact graph as reconstructing a coupled-fact graph: at each edit, a required fact comes from recent context or parametric memory, and the facts covered by neither form coherence debt.

Bardia Mohammadi, L. Klein, Aman Chadha et al. · 3 citations
#artificial intelligence Preprint Oct 2026

Teaching Agents to Code Reliably

Autonomous coding agents solve repository issues by reading code, running commands, editing files, and submitting patches. Extra inference-time compute yields gains only when it produces a useful repair and supplies reliable evidence for choosing one. Three behaviors decide both, and we argue they are teachable rather...

Muhammad Ahmed Mohsin, Myeongsoo Kim, Kang-Rui Ruan et al. · 0 citations
Preprint Oct 2026

Agent Skill Evolution: How Revisions Affect Coding Agents

Agent Skills, the SKILL.md files that tell an LLM coding agent how a project works, are revised like code, yet what a revision does to the agent is unknown. From 2,608 first/last revision pairs of 3,159 Skills, we characterize how Skills evolve and how they change together with the configuration of the agent's harness....

Jia-Jie Wang, Yu-Tong Zhao, Tian-Lin Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated re...

Guanqun Yang, Wen-Long Zhang, Tian Shi et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.