Skip to content
Conference

Lightweight Gated Residual Late Fusion for Frozen Vision Priors in Vision-Only Diffusion Policies

Aug 2026 · 2026 IEEE International Conference on Mechatronics and Automation (ICMA) · pp. 1216-1221 · 0 citations · 18 references

Abstract

Visuomotor diffusion policy has shown strong capability in robotic manipulation, yet its performance depends heavily on visual representation quality. Under vision-only settings, limited expressiveness of conventional encoders bottlenecks both policy performance and convergence. Large-scale self-supervised vision foundation models such as DINOv3 offer stronger visual priors, but how to stably integrate them without increasing system complexity remains underexplored.We propose a lightweight gated residual fusion method that adaptively injects features from a frozen vision foundation model into the policy input without modifying the diffusion backbone. A channel-wise gated residual mechanism dynamically modulates pretrained priors, improving visual conditioning while preserving the original policy architecture.Experiments on four robotic manipulation tasks in ManiSkill show consistent gains under the vision-only setting: compared with CNN-based diffusion policy, the average success rate improves from 0.59 to 0.74 (+25.4%), with training convergence accelerated by up to 57.7%; compared with fine-tuning methods, the training time is 5× lower with comparable average success. Ablation studies confirm that how pretrained features are integrated matters as much as features themselves. Gate-dynamic analysis further reveals an automatic early-to-late transition from pretrained priors to task-specific representations.These results demonstrate that vision foundation model priors can be effectively leveraged through a lightweight feature-fusion module without altering the policy backbone, providing a simple yet effective path toward stronger vision-only diffusion policies.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

QINA: Quantum-Inspired Nonlinear Adapters for Pretrained Vision Models

Adapting large pretrained vision models under limited data and frozen-backbone constraints remains a central challenge in transfer learning. While lightweight adapters and parameter-efficient fine-tuning methods are widely adopted, most rely on generic multilayer perceptrons or low-rank linear updates, offering limited...

M. M. Ghazi · 0 citations
Open access Aug 2026

Lightweight Redesign of Long-Used Operators in Vision Backbones for Efficient Visual Recognition

Recent vision backbones increasingly rely on sophisticated modules, whereas long-used operators such as residual connections, activations, and normalization layers remain less explored for lightweight redesign. This paper revisits these operators and proposes three operator-level redesigns: Subtractive Residual Connect...

Zhan-Yi Lian, Kepeng Luo, Yunfeng Wang · 0 citations
#machine learning Preprint Sep 2026

StarWM: Self-Supervised Trained Attention Routing for Robust World Models

StarWM is proposed, which uses a cross-attention module trained on self-supervised dynamics to decide where reconstruction applies and preserves state attributes with near-perfect fidelity through long-horizon imagination while systematically discarding distractors.

Zeqiang Zhang, Fabian Wurzberger, Maximilian Otte et al. · 0 citations
Preprint Sep 2026

Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control

Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned structure through large-scale image and video generation training. A growing line of robot policies builds on this generative prior, but how it should be transferred to control remains unclear, and existing appr...

Zan-Yi Wang, Yu-Heng Lei, Deng-Yang Jiang et al. · 0 citations
Aug 2026

Symbiotic Evolutionary Learning for Task-Adaptive Infrared and Visible Image Fusion.

This paper proposes a symbiotic evolutionary learning framework for task-adaptive IVIF, termed EvoFuse, which enables a mutual adaptation process where the fusion network and task models are jointly updated under a Pareto-inspired non-degradation criterion.

Jin-Yuan Liu, Bowei Zhang, Ludan Sun et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.