Skip to content
Conference

Raise One and Infer Three: Toward Reasoning- and Memory-Augmented Diffusion Policy Generalization

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · 0 citations · 37 references

TL;DR

This work proposes "raise one and infer three" diffusion policy (ROITDP), a novel approach that introduces two complementary mechanisms, including a reasoning mechanism built upon the Chain-of-Skill Noise Watermark, which enables temporally coherent multi-step reasoning throughout the diffusion process under distribution shifts.

Abstract

Diffusion policy has shown impressive performance in robotic manipulation tasks while struggling with out-of-distribution shifts and limited demonstrations. Recent advances primarily focus on improving geometric or perceptual representations for diffusion policy. However, these approaches rely heavily on instantaneous observations, making them vulnerable when visual inputs deviate from the training distribution. Our key insight is that robust generalization in diffusion policy requires both structured reasoning over skill compositions and an explicit memory mechanism for accumulating and reusing learned policy knowledge. To this end, we propose "raise one and infer three" diffusion policy (ROITDP), a novel approach that introduces two complementary mechanisms. First, we propose a reasoning mechanism built upon the Chain-of-Skill Noise Watermark, which encodes skill-level reasoning graph representations into the initial noise space, thereby enabling temporally coherent multi-step reasoning throughout the diffusion process under distribution shifts. We further introduce a memory mechanism in the form of a Self-evolving Policy Knowledge Bank that discretizes spatio-temporally refined trajectories into reusable skill primitives via a VQ-VAE architecture, providing adaptive memory to guide and refine action generation. Extensive experimental results on both simulated and real-world environments demonstrate the superiority and robustness of our method.

View source

Similar papers

Jul 2026

X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching

Pretraining navigation diffusion policies rely on large-scale expert demonstrations. These data are typically generated by a fully-informed oracle planner suited to a single nominal robot. This limits the policy's generalization to diverse embodiments and challenging scenarios (e.g., escaping dead ends or detouring lon...

Tian-Yu Yang, Yiming Zeng, Wenzhe Cai et al. · 0 citations
Preprint Sep 2026

DIA: Denoising Intermediate Advantage for Diffusion Policy Optimization

DIA is introduced, a policy-gradient method that learns a value function over partially denoised actions and uses it to construct a denoising level advantage for each step of the generative process, enabling it to discover more effective and efficient task-level strategies and subtask sequences that baseline methods fa...

Arjun Sohal, Yuchi Zhao, Miroslav Bogdanovic et al. · 0 citations
Preprint Jul 2026

AttriMem: Attribution-Guided Process Feedback for Agent Memory Learning

Experiments show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization, and outperforms retrieval-based, heuristic, and RL-based baselines on long-horizon dialogue question answering.

Qin-Feng Li, Yun-Tai Bao, Xinyang Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

PIVOT introduces a self-calibrated experience replay mechanism, which selectively collects and replays visually-grounded historical experiences as stable reference anchors for policy optimization, and designs a vision-guided advantage allocation mechanism to allocate additional vision-aware advantages to tokens based o...

Xin-Xin Song, Si-Yuan Li, Tingxiong Xiao et al. · 0 citations
Open access Aug 2026

Dual-Arm Manipulation Policy for Target-Cognitive Generalization

TCG-BP (Target-Cognitive Generalization Bimanual Policy), a target-prior-driven bimanual manipulation policy that converts language target descriptions into temporally consistent pixel-level target masks, and enhances visual representations through image–mask collaborative encoding and fusion is proposed.

Jianghao Sun, Pengjun Mao, LingJu Kong et al. · 0 citations
Preprint Aug 2026

From Memorization to Absorption: Mixed-Policy RL for Continual Knowledge Injection

This work proposes Golden-GRPO Injection (GRIN), a three-stage self-learning framework for continual knowledge injection that substantially outperforms SFT and mixed-policy RL baselines on the harder question types while matching them on basic fact recall.

Zhibo Hou, Fan Zhao, Zhiyu An et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.