VLN-AVP is proposed, a zero-shot navigation framework for AVP tasks that eliminates the dependency on pre-built maps, interprets semantic environmental contexts in parking scenarios, and enables intuitive navigation following natural language instructions and introduces a hybrid memory system.
Abstract
Existing methods in Autonomous Valet Parking (AVP) typically rely on pre-built maps, which severely restricts their scalability to unseen environments and open-vocabulary targets. Inspired by the application of Vision-Language Models (VLMs) in Vision-Language Navigation (VLN) tasks, we propose VLN-AVP, a zero-shot navigation framework for AVP tasks. By combining the precise spatial perception of a Bird's-Eye-View (BEV) model with the general intelligence of VLMs, our framework 1) eliminates the dependency on pre-built maps, 2) interprets semantic environmental contexts in parking scenarios, and 3) enables intuitive navigation following natural language instructions. Specifically, we introduce a hybrid memory system: a short-term perception memory tracks semantic visual cues to address the limitations of VLM's single-frame reasoning in existing methods, while a long-term topological memory facilitates stable policy learning from past experiences. To bridge the gap in existing benchmarks, we also present the VLN-AVP dataset and benchmark. Featuring 10 high-fidelity parking scenes and over 1,000 navigation episodes, it has the largest number of garage scenes to date and is the first VLN benchmark for underground parking. Extensive experiments demonstrate that in simulation, our method achieves an over 25% improvement in success rate compared to VLN methods and an over 15% improvement compared to other autonomous driving methods. Furthermore, it attains a leading success rate in real-world vehicle experiments, proving its practical feasibility.
Vision-and-Language Navigation (VLN) tasks require an agent to interpret natural language instructions and visual observations to navigate complex environments. Existing methods mostly construct topological or semantic maps and rely on the Large Language Model (LLM) for navigation decision-making, however, they still s...
Yuan Liu, Chuang Hu, Nan Ding· International Conferences on...· 0 citations
Continuous-environment vision-and-language navigation (VLN-CE) requires interpreting natural-language instructions in unseen 3D environments and executing continuous low-level actions. Existing methods often depend on LiDAR, panoramic cameras, or extra sensors; separate geometric-mapping and semantic-navigation visual...
Jian-He Zhao, Yan-Hua Qiu, Zhi-Yu Zhang et al.· 0 citations
HAM-VLN is presented, a decision-coupled, agent-authored memory that equips the robot with a persistent, depth-grounded world graph and reduces the context length by more than 65% compared to previous methods.
An Liu, Bingxi Liu, Hongyu Ding et al.· arXiv.org· 1 citation
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to navigate unseen environments by following natural language instructions. Current zero-shot VLN-CE methods either rely on pre-trained waypoint predictors or require multiple queries to large models per step. To address prohi...
Shi-Qi Pan, Qi Zheng, Hanmeng Sun et al.· 1 citation
Vision-Language Navigation (VLN) in unseen indoor environments is useful in real-world robotics, where an agent must follow natural-language instructions, locate objects, and answer spatial questions without a pre-built map or fixed object vocabulary. Multimodal vision-language models (VLMs) provide strong open-vocabul...
Long Giang Vu, Cheng-Kai Yao, Yu-Xin Liu et al.· 0 citations
LookStep is proposed, a unified end-to-end framework that combines Language Centric Future State Modeling and Event Driven Rolling Memory that uses language labels to generate coarse-grained navigation progress and future states for each candidate action, while autonomously deciding whether to write each observation in...
Kun-Yang Yu, Ying-Zhe Li, Hongyu Xu et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.