This work freezes the visual encoder and confine set propagation to a low-dimensional interface between it and the downstream policy, with the interface set calibrated from held-out camera-pose perturbations to achieve reachability analysis for visuomotor policies.
Abstract
Reachability analysis for visuomotor policies is difficult because large visual encoders make end-to-end set propagation computationally expensive and excessively conservative. We therefore freeze the visual encoder and confine set propagation to a low-dimensional interface between it and the downstream policy, with the interface set calibrated from held-out camera-pose perturbations. Propagating this set through the policy with zonotopes yields a terminal output-enclosure width that set-based training optimizes directly. During evaluation, camera-pose perturbations are sampled from the prescribed distribution, and rollout-level split conformal calibration converts the resulting action-deviation scores into a probabilistic reachable-action radius with finite-sample coverage. In controlled manipulation experiments, set-based training reduces this radius while preserving closed-loop task capability, and matched behavior-only, observational-consistency, and pointwise-adversarial controls all leave a larger radius.
Vision-Language-Action (VLA) models demonstrate strong generalization in robotic manipulation and navigation, but existing fine-tuning methods provide limited safety guarantees. Current approaches primarily rely on Lagrangian optimization that enforces safety through soft penalties on expected cumulative cost, often re...
Temporal Policy is introduced, a generative framework based on stochastic interpolants that formulates action generation as a temporally coupled transport problem and bypasses the computational bottleneck of independent Gaussian priors, helping enable high-frequency, closed-loop control.
Dylan Miller, Martin Jägersand· arXiv.org· 0 citations
This work introduces Robo-Dopamine 2.0, a history- and OOD-aware process reward model with a pairwise prediction interface that combines history-conditioned pairwise rewards that use source-aligned reference panels for synthetic OOD queries and observed rollout history for online queries, while preserving the queried e...
Yijie Xu, Hao-Peng Jin, Run Zhou et al.· 2 citations
Diffusion Policy has demonstrated strong performance in long-horizon robotic manipulation by generating smooth and executable action trajectories through conditional denoising. However, practical deployment remains limited by two key challenges. First, inadequate modeling of pose and position features for grasping and...
Zhe Wu, Yi-Cheng Shi, Yuan-Chong Wang et al.· 2026 IEEE International Conf...· 0 citations
Learning-based adversarial strategies can discover collision-inducing maneuvers, yet the generated behaviors become physically implausible when low-level execution relies on trajectory overwrites that bypass vehicle dynamics. This disconnect between the intended maneuver and its execution breaks kinematic consistency a...
Jian-Yu Zhao, Ruo-Han Zhao, Xuan Zhang et al.· IEEE/ASME International Conf...· 0 citations
World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-grained controllability with respect to low-level robot actions. A key obstacle to scaling such models in robotics is that actions are not a universal language in pixel space: changes in visual enviro...
Yuncong Yang, Zhen Han, Furkan Ozyurt et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.