May 2026· arXiv.org· Vol abs/2605.27955· 2 citations· 57 references
Computer Science
TL;DR
Skill-as-Pseudocode (SaP) is proposed, an automatic conversion of markdown skill libraries into typed pseudocode with deterministic quality control that wins paired games and falls below the prose baseline on the ALFWorld unseen split.
Abstract
Markdown skill libraries for LLM agents ship as free-form prose, forcing the agent to re-derive both the input schema and the concrete invocation syntax on every retrieval. This produces a"confused $\to$ re-retrieve $\to$ still confused"loop: the agent issues a partially-correct action, receives uninformative feedback, and re-retrieves the same prose. We propose Skill-as-Pseudocode (SaP), an automatic conversion of markdown skill libraries into typed pseudocode with deterministic quality control. From each cluster of similar procedural passages, SaP extracts a typed contract and filters it through a four-check deterministic verifier (coverage, binding, replacement, risk). Promoted contracts are inlined into a rewritten skill skeleton alongside restored action templates, giving the agent two complementary signals: a typed signature for what a skill does and a concrete template for how to invoke it. On the ALFWorld unseen split (134 games, gpt-4o-mini, three seeds), SaP wins 82/402 paired games versus 47/402 for the Graph-of-Skills (GoS) baseline (pooled McNemar $p = 8.2 \times 10^{-5}$), at $-22.8 \pm 6.4$% input tokens and $-14.5 \pm 4.1$% LLM calls per game. A bundle-component ablation attributes the gain to the pairing of typed contracts with concrete action templates: the contract alone falls below the prose baseline.
An AI agent's rebuild is only as good as the process that produced it. Prior work found that once a model is strong enough, a multi-agent rebuild pipeline loses to the simplest approach: giving the model the original code and one instruction (AgentModernize). We present rebuild-dossier, an open-source tool that locks a...
SKILLFORGE is introduced, a framework that decomposes formal code synthesis into a library of atomic, reusable skills, each targeting a specific subtask such as specification inference, body synthesis, invariant generation, error diagnosis, or targeted repair, and defined by a prompt template, tool binding, and decidab...
Yan-Ming Liu, Xinyue Peng, Jiannan Cao et al.· 0 citations
Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill s...
Yuxuan Liu, Hao-Ran Li, Yu-Hao Zhang et al.· 0 citations
SkillEffect is presented, a checked-lowering runtime for computations with a recoverable source relation, an audited bounded implementation, and a registered output postcondition that shows that one checked-lowering architecture can enforce heterogeneous registered memory relations at Agent tool dispatch.
Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing claims and invoking external search. Treating claims independently makes LLM and search calls scale with claim count and causes repeated searches for overlapping evidence a...
Ke-Ning Zheng, Ao-Ying Zheng, Zhi-Gang Chang et al.· 0 citations
Formal verification provides strong correctness guarantees, but its practical adoption is limited by the cost of writing precise formal specifications. While large language models can generate candidate specifications, prompt-only generation and monolithic LLM pipelines often struggle with verifier feedback, iterative...
Wen-Jie Wu, Jun-Jie Hu, Cheng Wen et al.· Companion Proceedings of the...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.