Compared with relying solely on initial observations and language instructions, predicting goal images with generative models as high-level visual guidance can significantly enhance the robustness of Vision-Language-Action (VLA) models. However, most existing foundation models have not systematically incorporated goal...
Xiao-Yuan Fang, Shuo Feng, Yuxuan Wang et al.· 0 citations
Experimental results on REVERIE and SOON datasets demonstrate that ViSMoE outperforms the previous state-of-the-art methods, showing the superiority of the proposed method.
Shuo Feng, Pi-Ji Li· International Conference on...· 0 citations
DeLS-Spec is proposed, a decoupled long-short context speculative decoding method that treats the fixed DFlash model as a long-context expert and introduces a lightweight local head as a short-context expert and consistently improves speedup and average acceptance length over DFlash across math, code, and dialogue benc...
AirAlign is proposed, a framework for RGB-only image-pair relative pose alignment for UAVs, using a pretrained visual geometry reconstruction model as the backbone to extract geometry-aware features from source-target image pairs.
Jin-Yi Zhou, Shuo Feng, Yufei Wu et al.· 0 citations
This work proposes MI-Distillation, a framework that constructs a continuous Instruct-Reasoning data spectrum through model interpolation, and introduces SeqLSS, which favors reasoning paths that are both informative and learnable for the student.
Yangsong Lan, Ren-Kai Hu, Hong-Kai Zheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.