Skip to content

Spotter: Let the Embodied Model Lead, and the VLM Reflect for It

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

Spotter is proposed, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control.

Abstract

Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in language models. We test whether they can repair a known error, which requires producing a correction and judging whether it is right. Stopped at a failure and allowed to retry, they seldom repair it through their own randomness or from a language description of the error, and best-of-N selection cannot pick the successful candidate after a failure. We attribute this to training only on successful demonstrations and to inputs too narrow to show what went wrong, and conclude that reflection must come from a vision-language model (VLM), which takes in far more information, such as the episode history and text, and is more general. Prior VLM-led work has the VLM plan every step and invoke the embodied model as a tool, placing the VLM on the critical path. We propose Spotter, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control. We run Spotter with Qwen and with GPT as the VLM, and both improve the embodied models; with GPT, Spotter improves Cosmos Policy and $\pi_{0.5}$ by 5.6 and 7.5 percentage points on RoboCasa, and raises $\pi_{0.5}$ from 47.2% to 57.0% on the Hard setting of RoboTwin 2.0 and from 53% to 83% on a real robot. Because the VLM steps in only when an error is confirmed, a successful episode with Qwen takes only 13 to 16 s longer than with the embodied model alone and about 70% less time than with a VLM-led baseline using the same model. Our code is available at https://github.com/zqc3117/Spotter.

View source

Similar papers

Preprint Sep 2026

Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

This work adds one feature-alignment term to ordinary VLA training: a frozen world model is run over the training frames once and cached, and the student learns to agree with that cache, indicating a broad representational prior rather than a fragile alignment between two particular networks.

T. Dao, Sankalp Yamsani, Jaden Park et al. · 0 citations
#machine learning Preprint Sep 2026

Part Grounding, Not Action Knowledge: Locating the Bottleneck in VLM Affordance Prediction

Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fails. We separate two steps that affordance questions conflate: identifying which part of an object to act on, and knowing what action that part requires. Across 19 articulat...

Sarthak Sattigeri · 0 citations
#artificial intelligence Preprint Oct 2026

Keeping JEPA World Models Plannable When Little of the Frame Moves

Specifying a goal in language rather than as a goal frame is a natural interface for planning with a latent world model, but testing it needs scenes in which language must discriminate between several objects. We build SLIM, a pushing benchmark with several small objects and paired visual and language goals on identica...

Florian Strohm, Patrick Wagner, Jannik Schwab et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DISEIL: Demonstration Distillation for Sample-Efficient Imitation Learning

A robot that can be taught a new task from a handful of demonstrations has to work out for itself what it still cannot do, and then ask for exactly that. Interactive imitation learning takes a step in that direction by letting a policy practice on its own and calling an expert when it goes wrong. Existing methods decid...

Suyog Khanal, V. ArunKumarA., Santu Rana · 1 citation
#artificial intelligence Preprint Aug 2026

Training-Free Action Correction for VLA Model Failures via Language Feedback

CorrectVLA is presented, a framework that translates task-level natural language corrections into additive action magnitude adjustments without modifying policy weights, and succeeds when policies possess strategic correctness and fails when fundamental comprehension is absent, establishing a practical operational boun...

O. Kwon, Pablo Ortega-Kral, A. Bucker et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.