The Role of Variability in Human Navigational Instructions in Visual Language Robot Navigation
Abstract
Visual Language Navigation (VLN) enables robots to follow natural language instructions to navigate visually perceived environments. Typically, VLN systems are trained on multi-modal datasets that pair visual scenes with navigation instructions. While prior work has focused on generalising to unseen environments, linguistic variability in human instructions remains largely unmodeled. This work explores the sensitivity of state-of-the-art VLN models to natural instructions that fall outside the highly constrained subsets found in standard benchmarks. We conduct a user study where diverse participants provide spoken navigational instructions to multiple robot embodiments. We then propose a high-level taxonomy to characterise variability of navigational instructions grounded in Human Robot Interaction (HRI). We use this framework to show that instruction style is influenced by human traits and robot embodiment. Finally, we evaluate multiple VLN model families on the collected data and demonstrate that altering instruction style alone, while keeping environments and trajectories fixed, leads to substantial performance degradation across all architectures. These findings highlight linguistic variability as a critical and underexplored challenge for robust, human-centered VLN systems.