Visually-Grounded Reward Synthesis (VGRS), which uses slow foundation models during training to produce fast robotic control policies, is proposed and achieves success rates above 55% on challenging long-horizon tasks while deploying successfully to real robots.
Abstract
Continuous robotic control requires policies that execute with low latency and modest computational cost during deployment. Foundation models provide strong semantic and visual reasoning, but repeatedly querying a large model throughout deployment incurs substantial inference latency and compute requirements. Language-to-Reward (L2R) methods avoid this deployment-time cost by using large language models (LLMs) to synthesize rewards for training lightweight policies, but these rewards are generated without visually analyzing how the learned policy physically fails, and thus often lack physical grounding. We propose Visually-Grounded Reward Synthesis (VGRS), which uses slow foundation models during training to produce fast robotic control policies. An LLM first synthesizes executable reward code from a natural-language instruction to train a lightweight hierarchical policy. When learning stalls, a frozen vision-language model (VLM) analyzes failed trajectories to provide failure mode diagnosis, which the LLM uses to rewrite and densify the reward. Since foundation models are used only during training, deployment requires only the learned policy. We perform experiments on simulated and real-world navigation and manipulation tasks, and show that VGRS achieves success rates above 55% on challenging long-horizon tasks while deploying successfully to real robots.
Learning long-horizon robot manipulation remains difficult and time-consuming, especially under sparse rewards due to inefficient exploration and reward assignment. We present a minimal integration of large language models (LLMs) with reinforcement learning (RL) in which the LLM is used strictly as an online action pro...
Meiyuan Gong, Yan Gao, Ze Ji· 2026 IEEE International Conf...· 0 citations
Quadruped robot locomotion policies are often trained using reinforcement learning, which in turn relies heavily on hand-crafted reward functions. Designing reward functions requires substantial manual engineering, and it is often unclear which local rewards will induce the desired global behavior. Shaped rewards from...
Merve Atasever, Keyan Azbijari, Cagan Bakirci et al.· 0 citations
This paper introduces a simple post-training recipe that turns off-the-shelf VLMs into robotic reasoners, and suggests that free-form language reasoning can function as a test-time compute mechanism for steering low-level policies.
Lehong Wu, Yuxiao Qu, Zhe-Yuan Hu et al.· 0 citations
This work instantiates Real-Time EXPO-FT, an RL framework for finetuning real-time VLA policies that meets the real-time control requirements of dynamic real-world manipulation, demonstrating rapid, sample-efficient adaptation to challenging real-world dynamics.
Perry Dong, Kuo-Han Hung, D. Sadigh et al.· 0 citations
The results demonstrate that the synthesized skill library enables the system to transfer to novel tasks with decreasing human intervention, providing a steerable and data-efficient alternative to black-box robot learning.
Daphne Chen, A. Jain, E. Goossen et al.· arXiv.org· 1 citation
Hierarchical Robotic Control (HiRoC) is proposed, a hierarchical post-training framework that decouples high-level task planning from low-level action execution and aligns the executor with planner-generated subgoals before reinforcement learning, mitigating the distribution misalignment between planning and execution.
He Kong, Ze Chen, Qi Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.