Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning

Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promising approach to improving robustness, yet existing methods commonly encode camera views independently and predict holistic scene dynamics without explicitly modeling their geometric relationships. We propose GWM-VLA, a geometry-aware latent world modeling framework for VLA learning. GWM-VLA combines geometry-aware multi-view state encoding, global context-conditioned target-view prediction, and shared latent-action representations grounded by robot-action supervision. Specifically, VGGT-$\Omega$ jointly aggregates multi-view observations at each timestep to construct geometry-aware multi-view states. The latent world model predicts the next-step patch tokens of a selected target view using patch and register tokens obtained after multi-view aggregation, thereby retaining multi-view geometric information without predicting the complete multi-view state. We use the wrist view as the target in our experiments, placing greater emphasis on end-effector motion and local gripper-object interactions. Finally, the shared latent-action representations condition both the latent world model and the flow-matching action head, allowing latent-prediction supervision and ground-truth robot-action supervision to jointly shape the same latent-action representations. Experiments across both simulation and real-world environments demonstrate the effectiveness and robustness of GWM-VLA.

Yanping Zhao, Hang Yu, Yiwei Wang et al. · 0 citations
Jul 2026

Self-Generating Reward Network for AUV Path Planning With Hybrid Global-Local Optimization.

Existing deep reinforcement learning (DRL) methods for autonomous underwater vehicle (AUV) path planning face two practical challenges: 1) dependency on manual reward engineering and 2) hyperparameter sensitivity in dynamic marine environments. This article presents a novel AUV path planning framework incorporating generative adversarial imitation learning (GAIL) and DRL algorithm that automates reward function synthesis through adversarial learning from expert demonstrations. The proposed architecture introduces a hierarchical reward mechanism that concurrently optimizes global trajectory planning and local motion constraints. By eliminating manual reward engineering, our approach reduces training complexity while maintaining policy convergence stability. Extensive experimentation demonstrates superior performance with 93.7% faster training convergence and 72.7% higher path convergence optimality compared to conventional DRL baselines. Two-tier validation confirms operational effectiveness: 1) Gazebo simulations achieve maximum 100% success rate in dynamic scenarios and 2) field deployments for submarine pipeline inspection attain 1 m average tracking accuracy. The results demonstrate that GAIL-DRL trained AUVs exhibit enhanced path planning stability while satisfying real-time planning requirements for marine transportation systems.

Chen Huang, Deshan Chen, Hao Feng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.