Skip to content
Preprint

"Dear LLaVA, Please Drive": A Depth-Aware Vision-Language Agent for Closed-Loop Robotic Control

Sep 2026 · 0 citations · 18 references
Computer Science

TL;DR

This work proposes a parameter-efficient approach to fine-tune a pretrained VLM for autonomous navigation using an Imperative Learning paradigm, and introduces a unified end-to-end navigation pipeline for natural-language-driven robotic control.

Abstract

Vision-language models (VLMs) provide a compelling foundation for reasoning-driven mobile navigation, offering rich contextual understanding and strong generalization from large-scale pretraining. Most existing navigation frameworks rely on imitation learning and therefore require substantial labeled trajectory data, limiting their scalability and robustness. In this work, we propose a parameter-efficient approach to fine-tune a pretrained VLM for autonomous navigation using an Imperative Learning paradigm. By optimizing against differentiable geometric cost fields rather than labeled trajectories, our model learns to generate collision-free paths exclusively from stereoscopic depth observations. We introduce a unified end-to-end navigation pipeline for natural-language-driven robotic control. This system leverages a shared VLM backbone with task-specific Low-Rank Adaptation (LoRA) modules, effectively bridging the gap from semantic target selection to low-level trajectory planning. Our approach achieves competitive Success weighted by Path Length (SPL) in unseen environments while updating less than 1% of the model's total parameters. Qualitative real-world experiments validate sim-to-real generalization and stable path planning without fine-tuning on real-world data. These results highlight a practical approach for deploying VLM-based agents on mobile robots, enabling high-level semantic navigation without the prohibitive requirement for large-scale, labeled trajectory data.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

TAO-DA: Towards Autonomous Operation--A Dual-Arm Vision-Language-Action Model for Coordinated Manipulation

A symmetric Dual-Arm Expert (DAE) architecture built upon a shared Vision-Language Model (VLM) backbone with decoupled, arm-specific expert towers is proposed, providing preliminary evidence of emergent skill generalization from single- to dual-arm tasks (as well as the reverse), together with cross-arm motion-domain s...

Yong-Shen Zhao, Han Gao, Bao-Ping Cheng et al. · 0 citations
Preprint Sep 2026

WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control

Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder,...

Timothy K. Johnsen, Marco Levorato · 0 citations
Preprint Sep 2026

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, rece...

Bing-Xuan Li, Si-Qi Song, Yi-Zhuo Wu et al. · 0 citations
Preprint Sep 2026

SAVLA: Symmetry-Aware Vision-Language-Action Models for Robotic Manipulation

Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the...

Jun-Le Li, Weixian Waylon Li, Fu-Xiang Wu et al. · 1 citation
Preprint Sep 2026

DistAL: Distance-based Advantage Learning for VLA Fine-Tuning

Vision-language-action models (VLAs) have trans- formed the field of robotic manipulation in recent years by combining the semantic understanding of LLMs with the precise control of flow-matching policies. Advantage conditioning is a recent technique that iteratively improves VLAs by training a value function on deploy...

Reece O'Mahoney, Ioannis Havoutis · 0 citations
Open access Aug 2026

SIRModel: Learning Spatial Intermediate Representation to Parameter-Efficiently Fine-Tune a Vision Language Model for Manipulation

Long-horizon robotic manipulation requires a policy to bridge task-level semantic reasoning with metric three-dimensional interaction geometry. Existing vision–language–action policies usually acquire geometry implicitly from visual tokens or introduce deterministic intermediate variables only in the image plane, which...

Li Lin, Ming-Hao Shi, Teng-Long Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.