Skip to content

VLN-AVP: Zero-Shot Vision-Language Navigation with Hybrid Long-Short-Term Memory for Autonomous Valet Parking

Jul 2026 · arXiv.org · Vol abs/2607.17767 · 0 citations · 32 references
Computer Science

TL;DR

VLN-AVP is proposed, a zero-shot navigation framework for AVP tasks that eliminates the dependency on pre-built maps, interprets semantic environmental contexts in parking scenarios, and enables intuitive navigation following natural language instructions and introduces a hybrid memory system.

Abstract

Existing methods in Autonomous Valet Parking (AVP) typically rely on pre-built maps, which severely restricts their scalability to unseen environments and open-vocabulary targets. Inspired by the application of Vision-Language Models (VLMs) in Vision-Language Navigation (VLN) tasks, we propose VLN-AVP, a zero-shot navigation framework for AVP tasks. By combining the precise spatial perception of a Bird's-Eye-View (BEV) model with the general intelligence of VLMs, our framework 1) eliminates the dependency on pre-built maps, 2) interprets semantic environmental contexts in parking scenarios, and 3) enables intuitive navigation following natural language instructions. Specifically, we introduce a hybrid memory system: a short-term perception memory tracks semantic visual cues to address the limitations of VLM's single-frame reasoning in existing methods, while a long-term topological memory facilitates stable policy learning from past experiences. To bridge the gap in existing benchmarks, we also present the VLN-AVP dataset and benchmark. Featuring 10 high-fidelity parking scenes and over 1,000 navigation episodes, it has the largest number of garage scenes to date and is the first VLN benchmark for underground parking. Extensive experiments demonstrate that in simulation, our method achieves an over 25% improvement in success rate compared to VLN methods and an over 15% improvement compared to other autonomous driving methods. Furthermore, it attains a leading success rate in real-world vehicle experiments, proving its practical feasibility.

View source

Similar papers

Conference Aug 2026

RO-VLMap: Real-Time Occupancy-Aware Visual Language Mapping for Robust Robot Navigation

Vision-and-Language Navigation (VLN) tasks require an agent to interpret natural language instructions and visual observations to navigate complex environments. Existing methods mostly construct topological or semantic maps and rely on the Large Language Model (LLM) for navigation decision-making, however, they still s...

Yuan Liu, Chuang Hu, Nan Ding · 0 citations
Preprint Sep 2026

LG-VLN: A Zero-Shot Vision-and-Language Navigation Framework with LangGraph State Orchestration

Continuous-environment vision-and-language navigation (VLN-CE) requires interpreting natural-language instructions in unseen 3D environments and executing continuous low-level actions. Existing methods often depend on LiDAR, panoramic cameras, or extra sensors; separate geometric-mapping and semantic-navigation visual...

Jian-He Zhao, Yan-Hua Qiu, Zhi-Yu Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

One MLLM, One Call: Efficient Zero-Shot Vision-and-Language Navigation via Spatial-Aware Waypoints

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to navigate unseen environments by following natural language instructions. Current zero-shot VLN-CE methods either rely on pre-trained waypoint predictors or require multiple queries to large models per step. To address prohi...

Shi-Qi Pan, Qi Zheng, Hanmeng Sun et al. · 1 citation
Preprint Sep 2026

AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation

Vision-Language Navigation (VLN) in unseen indoor environments is useful in real-world robotics, where an agent must follow natural-language instructions, locate objects, and answer spatial questions without a pre-built map or fixed object vocabulary. Multimodal vision-language models (VLMs) provide strong open-vocabul...

Long Giang Vu, Cheng-Kai Yao, Yu-Xin Liu et al. · 0 citations
Preprint Sep 2026

LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory

LookStep is proposed, a unified end-to-end framework that combines Language Centric Future State Modeling and Event Driven Rolling Memory that uses language labels to generate coarse-grained navigation progress and future states for each candidate action, while autonomously deciding whether to write each observation in...

Kun-Yang Yu, Ying-Zhe Li, Hongyu Xu et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.