Skip to content
Preprint

EndoWAM: A Grounded World-Action Model for Generalizable Endoscopic Navigation

Aug 2026 · 1 citation · 37 references
Computer Science

TL;DR

EndoWAM is presented, which is, to the authors' knowledge, the first WAM for generalizable robotic endoscopic navigation and introduces future grounding, which predicts task-relevant target regions in future observations from intermediate denoising features of a video world model.

Abstract

Autonomous endoscopic navigation can reduce clinicians'operational burden, yet robust control remains challenging due to tissue deformation, transient occlusions, and rapidly changing viewpoints. Existing learning-based policies typically predict actions from current observations without explicitly modeling future dynamics, limiting their robustness and reliability in safety-critical settings. World Action Models (WAMs) offer a promising alternative by coupling predictive visual dynamics with action generation, but extending them to robotic endoscopy remains challenging due to limited training data, restricted viewpoint diversity, deformable anatomy, and high inference latency. We present EndoWAM, which is, to our knowledge, the first WAM for generalizable robotic endoscopic navigation. EndoWAM introduces future grounding, which predicts task-relevant target regions in future observations from intermediate denoising features of a video world model. Specifically, EndoWAM couples a lightweight diffusion transformer for future target-region prediction with a discrete action expert through a shared predictive representation. This design injects target-aware supervision into predictive dynamics modeling, improving robustness to visual degradation and viewpoint changes while enabling real-time control in a single denoising pass. We further introduce EndoMotion, a robotic endoscopic motion dataset spanning three anatomically distinct procedures: ureteroscopy, esophagoscopy, and endoscopic retrograde cholangiopancreatography (ERCP). EndoWAM consistently outperforms all baselines and alternative grounding strategies, while demonstrating strong zero-shot generalization to unseen viewpoints, environments, and targets. These results establish EndoWAM as a predictive, target-grounded framework for accurate, generalizable, and long-horizon navigation in visually constrained endoscopic environments.

View source

Similar papers

Preprint Aug 2026

GaussianDream++: Efficient 3D Gaussian World Modeling for Robotic Manipulation

GaussianDream demonstrates that training-time current Gaussian reconstruction and future Gaussian prediction provide effective 3D supervision, but its dense VGGT/TGE-based prefix jointly carries state, dynamics, and action-conditioning information.

Yu-Qing Jiang, Zi-Jian Zhang, Wei-Tao Zhou et al. · 0 citations
Preprint Aug 2026

GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

GeniWorld is presented, an interactive world model for robots that generalizes robustly across unseen scenarios by explicitly decoupling embodiment kinematics from environmental dynamics, and generates diverse manipulation trajectories within the world model, improving downstream policy performance and robustness in co...

Cheng-Hao Gu, Hanyang Yu, Jingbo Zhang et al. · 2 citations · ⚡2
Jul 2026

Action-Conditioned World Model for Goal Plane Probe Guidance in Robotic Ultrasound

We present an action-conditioned world model framework for goal plane probe guidance in robotic ultrasound, with a focus on neck ultrasound scanning. Autonomous ultrasound tasks often require large numbers of probe-motion trajectories for training, but collecting high-quality demonstrations is labor-intensive and expli...

Si-Qi Fan, Mingcong Chen, Rang Liu et al. · 0 citations
Preprint Aug 2026

LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models

LAWM-3D is proposed, which introduces three tightly coupled key designs: a multi-view invariant unified action tokenization scheme for learning 3D-aware latent actions, a geometric alignment constraint that anchors intermediate encoder features to a pretrained 3D foundation model, and a non-injective RGB-D joint recons...

Jia-Rui Yang, Jiale Zhange, Jiawei Li et al. · 1 citation
Preprint Aug 2026

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

SimWAM, a simple yet effective WAM that leverages future-video prediction as a training-time supervision signal, co-trains a pretrained video expert and a lightweight action expert with joint flow matching and applies reinforcement learning to optimize a compositional driving reward beyond trajectory imitation.

Zongchuang Zhao, Xin Zhou, Tianyang Xu et al. · 4 citations
Review Aug 2026

Surgical Video Generation From Diffusion to World Models: A Survey

Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition remains constrained by privacy, cost, and class imbalance. Surgical video generation has emerged as a transformative approa...

Fuxiang Huang, Chenxu Zhang, Liang Han et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.