Images capture the world at one moment, whereas video reveals how it changes. Image self-supervision learns spatial structure from a single moment, while video methods commonly learn temporal relationships inside a representation computed jointly from several frames. We ask whether watching a scene change can instead i...
Wen Huang, Hang Guo, Jia-Rui Yang et al.· 0 citations
Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation, yet complex contact-rich tasks often benefit from multi-camera observations that jointly capture the end effector, objects, and targets under occlusion. Existing multi-camera VLAs usually concatenate view tokens, leaving actio...
Jia-Rui Yang, Ye-Hao Lu, Yu-Ning Su et al.· 1 citation
Online post-training of vision-language-action (VLA) models requires efficient use of robot interaction and reliable policy improvement from continually collected experience. We propose asynchronous Replay-Anchored Policy improvement (RAPolicy), a framework that performs rollout and learning concurrently while groundin...
Jia-Rui Yang, Jia-Jin Zhang, Bin Zhu et al.· 0 citations
LAWM-3D is proposed, which introduces three tightly coupled key designs: a multi-view invariant unified action tokenization scheme for learning 3D-aware latent actions, a geometric alignment constraint that anchors intermediate encoder features to a pretrained 3D foundation model, and a non-injective RGB-D joint recons...
Jia-Rui Yang, Jiale Zhange, Jiawei Li et al.· 1 citation
Across three flow-based VLA models on multiple simulated manipulation benchmarks and two real-world tasks, StructRL improves exploration efficiency and OOD performance over prior in-chain baselines, demonstrating the effectiveness of structured action-space exploration for adapting flow-based VLA with RL.
Jia-Rui Yang, Bin Zhu, Jing-Jing Chen et al.· 1 citation
This paper argues that what a VLA needs is not the ability to generate language, but the ability to consume grounded language, and introduces a framework that endows a VLA with language competence through in-context post-training and an agentic tool-use interface.
Jia-Rui Yang, Wen Huang, Jia-Le Zhang et al.· 1 citation
V-Link is proposed, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer and injects them into Action DiT through asymmetric pathways.
Ye-Hao Lu, Jia-Rui Yang, Yu-Ning Su et al.· 0 citations
TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment, and shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation.
Jia-Rui Yang, Ye-Hao Lu, Yu-Ning Su et al.· 1 citation
OVIP-SG is presented, a unified framework for instance-preserving semantic mapping, functional scene partitioning, and language-guided small, fine-grained object retrieval that outperforms ConceptGraphs under a unified evaluation protocol on Replica.
Tianjing Hao, Hai-Yu Lan, Ang Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.