Skip to content
Conference Open access

ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding

Aug 2026 · International Conference on Artificial Neural Networks · pp. 314-325 · 0 citations · 24 references
Computer Science

TL;DR

Experimental results on REVERIE and SOON datasets demonstrate that ViSMoE outperforms the previous state-of-the-art methods, showing the superiority of the proposed method.

Abstract

Embodied Referring Expression Grounding is the task of enabling an agent to navigate in real environments and to localize a remote object based on natural language instructions. In this scenario, the agent needs to select one view for navigation at each step and identify a specific object among all candidate objects at the destination. However, most of the previous approaches fail to distinguish between views and objects, instead processing them using the vanilla vision encoder, which results in ambiguous representations of both views and objects. To address the above issues, we propose ViSMoE, which equips sparse Mixture-of-Experts with a visual-aware routing policy for the embodied agent. This framework processes different types of visual information specifically, resulting in discriminative visual representations for both views and objects. Experimental results on REVERIE and SOON datasets demonstrate that ViSMoE outperforms the previous state-of-the-art methods, showing the superiority of our proposed method.

Read PDF

Similar papers

#artificial intelligence Preprint Aug 2026

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

LightNav-0 is presented, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads, and establishes compact VLMs as a unified and transferable backbone for generalist embodied navigation.

Shao-An Wang, Ao-Cheng Luo, Fei Huang et al. · 2 citations
Conference Aug 2026

Semantic relevance guided grounding for MLLM-based embodied navigation

Experiments on the AI2Thor platform demonstrate that SRG-Nav outperforms baseline methods in both success rate and path efficiency, validating that structured semantic-visual prompts significantly improve the robustness of embodied navigation.

Shuai Chen, Hao Chen, Beiyu Guo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents

Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce...

Rui-Xun Liu, Yuxuan Wang, Jia-Cheng Xie et al. · 1 citation · ⚡1
Preprint Sep 2026

From Gaze to Meaning: A Training-Free AI Agent for Unified Grounding and Explanation

This work introduces the first training-free Gaze Target Agent (GTA) for gaze-guided reasoning across tasks such as gaze target prediction, attention localization, and object identification by leveraging pretrained vision-language models, augmenting them with visually guided prompts, and employing a memory-based retrie...

Shayan Nasiriboukani, Sara Atito, Mohammad Nezamipour et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EviRover: Reinforcing Agentic Perception Beyond a Glance

Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or requi...

Kai-Xuan Fan, Kai-Tuo Feng, Tian-Shuo Peng et al. · 0 citations
Preprint Aug 2026

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

It is shown that MVRD makes visual representations more geometric while retaining language alignment, and generalizes to 3D scene understanding tasks such as object grounding, dense captioning, and question answering, while approaching feature fusion methods with considerably fewer added parameters and lower latency.

K. T. Nguyen, Hanbo Shim, Jinwoo Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.