Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee com...
Bo Mao, Hang He, Lin-Ting Wang et al.· 0 citations
TSGen, an automated pipeline for generating high-quality, structured TSGs from historical incident reports using large language models (LLMs), consists of filtering and classifying incident data into diagnostically relevant categories and distilling core incidents to ensure diversity and generalizability.
Yi Xiao, Hongyu Zhang, Dr. I. I. Genkin et al.· SIGSOFT FSE Companion· 0 citations
This work introduces WebXSkill, a framework that bridges a grounding gap with executable skills, each pairing a parameterized action program with step-level natural-language guidance, and finds that better skill deployment mode depends on a model's plan and execution capability.