A contract-grounded BT synthesis architecture in which a coding agent queries a robot-side Model Context Protocol (MCP) server to retrieve an explicit contract consisting of a skill library, permitted BT operators, and optional BT composition templates, before synthesizing a BT for validation and execution is proposed.
Abstract
Synthesizing deployable robot behavior trees (BTs) from natural language (NL) requires grounding to ensure every generated BT references only skills a robot can actually execute. Existing LLM-based BT synthesis approaches often place this grounding responsibility on the prompt author. This makes deployment brittle when the author does not know which skills the robot can execute, how those skills are parameterized, or how the robot runtime software constrains valid BT structure. This paper proposes a contract-grounded BT synthesis architecture in which a coding agent queries a robot-side Model Context Protocol (MCP) server to retrieve an explicit contract consisting of a skill library, permitted BT operators, and optional BT composition templates, before synthesizing a BT for validation and execution. In our framework, non-expert operators issue NL commands without knowledge of robot implementation details, while a robot runtime validation gate enforces correctness before execution. We evaluate two LLMs, a closed model (Sonnet 4.6) and a smaller open-source model (Gemma4:31b), across 110 simulated tasks in PyRoboSim and 14 tasks on a physical Husarion Panther robot. Results show that contract grounding enables near-perfect BT validation and high task success, that BT composition templates substantially recover success on reactive control-flow tasks for the smaller model, and that the architecture transfers to physical hardware running a Nav2 stack opaque to both operator and agent.
This work formalizes the reasoning-execution boundary as a typed contract and constrains language-level decisions through schema-validated tool calls defined by the Model Context Protocol, rejecting malformed commands before they reach the robot.
CEDAR is presented, a counterexample-guided framework that grounds instructions as regular languages over environment event traces and represents both skills and specifications as deterministic finite automata, suggesting that regular languages offer a practical verification layer between natural-language instructions...
Le Chen, Alvaro Velasquez, Ashutosh Trivedi· 0 citations
Learning from Demonstration (LfD) combined with Behavior Trees (BTs) aims to lower the programming burden required to construct robot task programs from demonstrations. However, existing approaches have two structural limitations: skills are bound to specific object instances with no mechanism for runtime rebinding, an...
Bo-Gang Jiang, Peng-Ji Wu, Zhi-Jie Xu et al.· IEEE Robotics and Automation...· 0 citations
Translating ambiguous human instructions into verifiable robot actions in semistructured environments requires bridging the semantic–execution gap between high-level intent and formally monitored execution. Existing symbolic planners rely on fixed, manually specified state vocabularies, while learning-based methods off...
Zhang-Li Zhou, Chang-Mao Chen, Hao Li et al.· IEEE Transactions on robotic...· 0 citations
Overall, Code2Skill transforms procedural knowledge embedded in repositories into grounded, verifiable, and transferable agent skills, providing initial evidence that the pipeline can expand with the growing volume of AI-generated software.
Yong-Qi Tong, Pan Wang, Hang Wang et al.· 1 citation
ProtoAct, a structured protocol-grounding framework that converts free-form biological procedures into state-aware, embodiment-ready action sequences, is presented, providing a practical interface between biological protocol understanding and embodied robotic execution.
Zhe Liu, Jiaming Gu, Zhao-Hui Du et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.