Aug 2026· 2026 IEEE International Conference on Mechatronics and Automation (ICMA)· pp. 1216-1221· 0 citations· 18 references
Abstract
Visuomotor diffusion policy has shown strong capability in robotic manipulation, yet its performance depends heavily on visual representation quality. Under vision-only settings, limited expressiveness of conventional encoders bottlenecks both policy performance and convergence. Large-scale self-supervised vision foundation models such as DINOv3 offer stronger visual priors, but how to stably integrate them without increasing system complexity remains underexplored.We propose a lightweight gated residual fusion method that adaptively injects features from a frozen vision foundation model into the policy input without modifying the diffusion backbone. A channel-wise gated residual mechanism dynamically modulates pretrained priors, improving visual conditioning while preserving the original policy architecture.Experiments on four robotic manipulation tasks in ManiSkill show consistent gains under the vision-only setting: compared with CNN-based diffusion policy, the average success rate improves from 0.59 to 0.74 (+25.4%), with training convergence accelerated by up to 57.7%; compared with fine-tuning methods, the training time is 5× lower with comparable average success. Ablation studies confirm that how pretrained features are integrated matters as much as features themselves. Gate-dynamic analysis further reveals an automatic early-to-late transition from pretrained priors to task-specific representations.These results demonstrate that vision foundation model priors can be effectively leveraged through a lightweight feature-fusion module without altering the policy backbone, providing a simple yet effective path toward stronger vision-only diffusion policies.
Adapting large pretrained vision models under limited data and frozen-backbone constraints remains a central challenge in transfer learning. While lightweight adapters and parameter-efficient fine-tuning methods are widely adopted, most rely on generic multilayer perceptrons or low-rank linear updates, offering limited...
Experimental results demonstrate that adaptability is not solely determined by model size, but rather by how effectively parameter plasticity is regulated in dynamic environments.
Xiao-Rong Zeng, Weiqiang Chen, Peng Shi et al.· 0 citations
Recent vision backbones increasingly rely on sophisticated modules, whereas long-used operators such as residual connections, activations, and normalization layers remain less explored for lightweight redesign. This paper revisits these operators and proposes three operator-level redesigns: Subtractive Residual Connect...
StarWM is proposed, which uses a cross-attention module trained on self-supervised dynamics to decide where reconstruction applies and preserves state attributes with near-perfect fidelity through long-horizon imagination while systematically discarding distractors.
Zeqiang Zhang, Fabian Wurzberger, Maximilian Otte et al.· 0 citations
Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned structure through large-scale image and video generation training. A growing line of robot policies builds on this generative prior, but how it should be transferred to control remains unclear, and existing appr...
Zan-Yi Wang, Yu-Heng Lei, Deng-Yang Jiang et al.· 0 citations
This paper proposes a symbiotic evolutionary learning framework for task-adaptive IVIF, termed EvoFuse, which enables a mutual adaptation process where the fusion network and task models are jointly updated under a Pareto-inspired non-degradation criterion.
Jin-Yuan Liu, Bowei Zhang, Ludan Sun et al.· IEEE Transactions on Pattern...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.