Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· 0 citations· 37 references
TL;DR
This work proposes "raise one and infer three" diffusion policy (ROITDP), a novel approach that introduces two complementary mechanisms, including a reasoning mechanism built upon the Chain-of-Skill Noise Watermark, which enables temporally coherent multi-step reasoning throughout the diffusion process under distribution shifts.
Abstract
Diffusion policy has shown impressive performance in robotic manipulation tasks while struggling with out-of-distribution shifts and limited demonstrations. Recent advances primarily focus on improving geometric or perceptual representations for diffusion policy. However, these approaches rely heavily on instantaneous observations, making them vulnerable when visual inputs deviate from the training distribution. Our key insight is that robust generalization in diffusion policy requires both structured reasoning over skill compositions and an explicit memory mechanism for accumulating and reusing learned policy knowledge. To this end, we propose "raise one and infer three" diffusion policy (ROITDP), a novel approach that introduces two complementary mechanisms. First, we propose a reasoning mechanism built upon the Chain-of-Skill Noise Watermark, which encodes skill-level reasoning graph representations into the initial noise space, thereby enabling temporally coherent multi-step reasoning throughout the diffusion process under distribution shifts. We further introduce a memory mechanism in the form of a Self-evolving Policy Knowledge Bank that discretizes spatio-temporally refined trajectories into reusable skill primitives via a VQ-VAE architecture, providing adaptive memory to guide and refine action generation. Extensive experimental results on both simulated and real-world environments demonstrate the superiority and robustness of our method.
Pretraining navigation diffusion policies rely on large-scale expert demonstrations. These data are typically generated by a fully-informed oracle planner suited to a single nominal robot. This limits the policy's generalization to diverse embodiments and challenging scenarios (e.g., escaping dead ends or detouring lon...
Tian-Yu Yang, Yiming Zeng, Wenzhe Cai et al.· arXiv.org· 0 citations
DIA is introduced, a policy-gradient method that learns a value function over partially denoised actions and uses it to construct a denoising level advantage for each step of the generative process, enabling it to discover more effective and efficient task-level strategies and subtask sequences that baseline methods fa...
Arjun Sohal, Yuchi Zhao, Miroslav Bogdanovic et al.· 0 citations
Experiments show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization, and outperforms retrieval-based, heuristic, and RL-based baselines on long-horizon dialogue question answering.
Qin-Feng Li, Yun-Tai Bao, Xinyang Yu et al.· 0 citations
PIVOT introduces a self-calibrated experience replay mechanism, which selectively collects and replays visually-grounded historical experiences as stable reference anchors for policy optimization, and designs a vision-guided advantage allocation mechanism to allocate additional vision-aware advantages to tokens based o...
Xin-Xin Song, Si-Yuan Li, Tingxiong Xiao et al.· 0 citations
TCG-BP (Target-Cognitive Generalization Bimanual Policy), a target-prior-driven bimanual manipulation policy that converts language target descriptions into temporally consistent pixel-level target masks, and enhances visual representations through image–mask collaborative encoding and fusion is proposed.
Jianghao Sun, Pengjun Mao, LingJu Kong et al.· Electronics· 0 citations
This work proposes Golden-GRPO Injection (GRIN), a three-stage self-learning framework for continual knowledge injection that substantially outperforms SFT and mixed-policy RL baselines on the harder question types while matching them on basic fact recall.
Zhibo Hou, Fan Zhao, Zhiyu An et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.