Gradient-free Robot Action generation via Combined diffusion-MPPI posterior mean Estimation (GRACE), which guides a pretrained diffusion policy with Model Predictive Path Integral (MPPI) control using only forward cost evaluations.
Abstract
Diffusion policies generate multimodal robot action sequences from demonstrations, but steering them toward deployment-time constraints typically relies on differentiable guidance costs. This excludes many practical safety constraints, such as binary collision checks, joint limits, and black-box rollout costs that are nondifferentiable. We propose Gradient-free Robot Action generation via Combined diffusion-MPPI posterior mean Estimation (GRACE), which guides a pretrained diffusion policy with Model Predictive Path Integral (MPPI) control using only forward cost evaluations. Building on the common score-ascent structure of diffusion and MPPI, GRACE constructs a cost-conditioned guidance posterior at each reverse step and estimates its mean with a single MPPI update centered at the diffusion reverse mean. For differentiable costs, GRACE recovers conventional gradient guidance under a first-order, matched-covariance approximation. GRACE attains higher success rates than diffusion-based and sampling-based baselines in simulation. On a real 7-DoF manipulator, GRACE avoids a deployment-time obstacle that the unguided prior collides with in every trial. Code and experiment videos are available at https://anonymous.4open.science/w/grace-70BB/.
Diffusion policies have demonstrated excellent performance in robotic control tasks, yet their reliance on 50 to 100 denoising steps impedes real-time deployment. Existing acceleration methods based on sampler improvements lack responsiveness to perceptual quality and cannot adaptively adjust computation. Moreover, sta...
Qi Chen, Xinyang Ren, Jiajun Xing et al.· IEEE Robotics and Automation...· 0 citations
A data-efficient diffusion RL post-training framework - GQRM (Group Q-score Reweighted Matching), which achieves state-of-the-art cross-embodiment visual navigation performance and introduces two complementary designs: a self-bootstrapped exploration strategy with behavior perturbation that preserves the pretrained pol...
Tian-Yu Yang, Yi-Ming Zeng, Wen-Zhe Cai et al.· arXiv.org· 0 citations
Temporal Policy is introduced, a generative framework based on stochastic interpolants that formulates action generation as a temporally coupled transport problem and bypasses the computational bottleneck of independent Gaussian priors, helping enable high-frequency, closed-loop control.
Dylan Miller, Martin Jägersand· arXiv.org· 0 citations
Decentralized multi-robot navigation is difficult when robots must act from local observations without centralized coordination or explicit inter-robot communication. A belief-driven hybrid reinforcement learning framework is evaluated for planar multi-robot navigation under partial observability. Each robot builds a c...
V. Malathi, Pramod Sreedharan, Rthuraj Puthiyaveedu Rajesh et al.· Robotics· 0 citations
Planning Diffusion Policy Optimization is proposed, an offline-to-online reinforcement-learning framework that uses a diffusion policy to generate short-horizon action chunks for crowd navigation and obtains an improved success rate over strong baselines and ablations demonstrate that action chunks are especially impor...
A structured scoping review of reinforcement learning for generative robot policies and a bounded state-based locomotion reproduction is presented and a full-chain backpropagation adaptation exhibited clear seed-dependent variation.
Shihan Sun, Yinlong Liu· Robotics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.