Skip to content

Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents

May 2026 · arXiv.org · Vol abs/2605.27955 · 2 citations · 57 references
Computer Science

TL;DR

Skill-as-Pseudocode (SaP) is proposed, an automatic conversion of markdown skill libraries into typed pseudocode with deterministic quality control that wins paired games and falls below the prose baseline on the ALFWorld unseen split.

Abstract

Markdown skill libraries for LLM agents ship as free-form prose, forcing the agent to re-derive both the input schema and the concrete invocation syntax on every retrieval. This produces a"confused $\to$ re-retrieve $\to$ still confused"loop: the agent issues a partially-correct action, receives uninformative feedback, and re-retrieves the same prose. We propose Skill-as-Pseudocode (SaP), an automatic conversion of markdown skill libraries into typed pseudocode with deterministic quality control. From each cluster of similar procedural passages, SaP extracts a typed contract and filters it through a four-check deterministic verifier (coverage, binding, replacement, risk). Promoted contracts are inlined into a rewritten skill skeleton alongside restored action templates, giving the agent two complementary signals: a typed signature for what a skill does and a concrete template for how to invoke it. On the ALFWorld unseen split (134 games, gpt-4o-mini, three seeds), SaP wins 82/402 paired games versus 47/402 for the Graph-of-Skills (GoS) baseline (pooled McNemar $p = 8.2 \times 10^{-5}$), at $-22.8 \pm 6.4$% input tokens and $-14.5 \pm 4.1$% LLM calls per game. A bundle-component ablation attributes the gain to the pairing of typed contracts with concrete action templates: the contract alone falls below the prose baseline.

View source

Similar papers

Preprint Aug 2026

Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal

An AI agent's rebuild is only as good as the process that produced it. Prior work found that once a model is strong enough, a multi-agent rebuild pipeline loses to the simplest approach: giving the model the original code and one instruction (AgentModernize). We present rebuild-dossier, an open-source tool that locks a...

P. Fawcett · 0 citations
#natural language process... Preprint Aug 2026

SkillForge: Compositional Skill Synthesis with Verification-in-the-Loop for Generating Formally Verified Dafny Programs

SKILLFORGE is introduced, a framework that decomposes formal code synthesis into a library of atomic, reusable skills, each targeting a specific subtask such as specification inference, body synthesis, invariant generation, error diagnosis, or targeted repair, and defined by a prompt template, tool binding, and decidab...

Yan-Ming Liu, Xinyue Peng, Jiannan Cao et al. · 0 citations
#artificial intelligence Review Oct 2026

SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown

Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill s...

Yuxuan Liu, Hao-Ran Li, Yu-Hao Zhang et al. · 0 citations
Preprint Aug 2026

SkillEffect: Checked Lowering for Memory-Bounded Agent Tools

SkillEffect is presented, a checked-lowering runtime for computations with a recoverable source relation, an audited bounded implementation, and a registered output postcondition that shows that one checked-lowering architecture can enforce heterogeneous registered memory relations at Agent tool dispatch.

Yi-Nuo Wang, Yiyu Shi · 0 citations
#machine learning Preprint Sep 2026

EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy

Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing claims and invoking external search. Treating claims independently makes LLM and search calls scale with claim count and causes repeated searches for overlapping evidence a...

Ke-Ning Zheng, Ao-Ying Zheng, Zhi-Gang Chang et al. · 0 citations
Book Open access Oct 2026

Spec-Skill: A Pluggable Coding-Agent Plugin for Neuro-symbolic Program Specification Synthesis

Formal verification provides strong correctness guarantees, but its practical adoption is limited by the cost of writing precise formal specifications. While large language models can generate candidate specifications, prompt-only generation and monolithic LLM pipelines often struggle with verifier feedback, iterative...

Wen-Jie Wu, Jun-Jie Hu, Cheng Wen et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.