Jul 2026
SoftNav: Injecting 3D Scene Tokens into VLMs for Embodied Navigation
SoftNav is introduced, which injects entity-level 3D continuous representations -- one token per detected object or frontier -- into a VLM's hidden space as soft tokens through a lightweight projector, enabling transferable navigation with minimal training.
Yi Wu, Junjie An, Xiao Liu et al.
· arXiv.org · 0 citations