This work introduces BioVLN, a simulation platform for developing and evaluating visual-language navigation agents in biomedical laboratories and shows that geometric exploration reaches 74.4--87.5% success, while sampling multiple valid positions in the operation area improves success and reduces unsafe proximity.
Zhe Liu, Quan Lu, Zhao-Hui Du et al.· arXiv.org· 0 citations
LightNav-0 is presented, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads, and establishes compact VLMs as a unified and transferable backbone for generalist embodied navigation.
Shao-An Wang, Ao-Cheng Luo, Fei Huang et al.· 3 citations
This paper presents the first systematic benchmark of 17 edge-deployable SLMs against 4 online APIs for robotic navigation instruction decomposition, and proposes a lightweight hybrid semantic-geometric goal localization framework that combines open-vocabulary object detection, prompted segmentation, and LiDAR geometry...
Ali Salmasi, Xian-Jia Yu, Tomi Westerlund· arXiv.org· 0 citations
Visual Language Navigation (VLN) enables robots to follow natural language instructions to navigate visually perceived environments. Typically, VLN systems are trained on multi-modal datasets that pair visual scenes with navigation instructions. While prior work has focused on generalising to unseen environments, lingu...
Malak Sayour, Pamela Carreno-Medrano, Michael Burke et al.· IEEE Robotics and Automation...· 0 citations
This work proposes a user-friendly, interactive framework that eliminates the reliance on globally consistent maps and integrates visual perception with Large Language Models (LLM) to interpret user commands via text or voice.
Praveen Kumar, K. Guruprasad, Tushar Sandhan· 0 citations
Experimental results in various task scenarios show that the proposed framework consistently improves overall task success rates compared with unimodal settings with different LLMs and achieves a higher success rate compared to using only visual or force data.