Teleoperated demonstrations are a primary source of data for robot manipulation, and teleoperated interventions are a primary mechanism for correcting policies at deployment. Yet most teleoperation systems close the loop through vision alone and are built around parallel-jaw grippers, limiting both what the robot can e...
Zhan-Peng He, Joaquin Palacios, Zhang-Yu Wang et al.· 0 citations
Indoor 3D referring segmentation aims to identify and segment the target object in a point cloud according to a natural language expression. Although recent advances in multimodal feature fusion have markedly improved this task, existing methods do not jointly model the structural regularities commonly observed in indo...
Li Yuan, Bo Kong, Chenhao Li et al.· Symmetry· 0 citations
Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Models (VLMs) as reward models, computing text-observation similarity to bypass manual reward engineering. Although promising, these rewards are often noisy and unreliable, l...
Pyrros Koussios, Chen-Hao Li, Xin Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.