Skip to content
Preprint

VLA-Corrector: Stage-Aware Observable State Understanding for Prompt-Based Closed-Loop Recovery of Vision-Language-Action Policies

Sep 2026 · 0 citations · 51 references
Computer Science

TL;DR

Results show that observable visual-proprioceptive-action history is sufficient to infer latent task states and enable practical failure recovery for existing VLA policies, and the proposed approach substantially improves closed-loop reliability under diverse perturbations.

Abstract

Long-horizon robot manipulation with Vision-Language-Action (VLA) policies remains vulnerable to execution-time deviations, as final task success provides little information for diagnosing and correcting failures caused by action noise, object displacement, or goal misalignment. We introduce a stage-aware failure verification and Prompt Recovery framework that enables closed-loop correction of a fixed VLA policy without parameter updates or privileged simulator states. The framework introduces an observable-history-based Learned Verifier that jointly estimates manipulation progress and execution risk by temporally modeling multi-view visual observations, proprioceptive states, and executed actions. To provide interpretable task understanding, we represent manipulation execution through semantic progress stages, including approach, alignment, grasp, transport, and placement, and identify stage-specific failure patterns. Upon detecting abnormal execution, the framework preserves the original instruction and generates a stage-conditioned recovery prompt, allowing the same frozen VLA policy to produce corrective actions. Extensive multi-round evaluations on LIBERO and LIBERO Plus demonstrate that the proposed approach substantially improves closed-loop reliability under diverse perturbations. Without access to privileged object or goal coordinates, the Learned Verifier achieves recovery performance close to that of the privileged rule-based verifier in the evaluated settings. These results show that observable visual-proprioceptive-action history is sufficient to infer latent task states and enable practical failure recovery for existing VLA policies.

View source

Similar papers

Preprint Sep 2026

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies

This work proposes CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution, and introduces the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anoma...

Jun-Lan Xiao, Jun-Wei Jiang, Zai-Bin Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

AGM: Achievement-Grounded Memory for Closed-Loop Agents with Frozen VLA Policies

Achieve-Grounded Memory is proposed, a lightweight closed-loop framework for frozen VLA policies that represents a task as a subgoal sequence with a progress pointer and advances this memory only after the current subgoal is verified by physical evidence.

Hong-Bo Gao, Ze-Yu Ni, Xin Wen et al. · 1 citation
Preprint Sep 2026

ProAct-VLM: Pre-Failure Vision-Language Task Replanning with Continuous Perception Feedback

Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly. Traditional rule-based decision-making pipelines are brittle...

Ahmed N. Ahmed, Omar Moured, Mughni Irfan Mohammed Abdul et al. · 0 citations
Preprint Aug 2026

Imagining Recovery: Inference-Time Counterfactual Realignment for Vision-Language-Action Models

Counterfactual Realignment (CoRe), a training-free framework that recovers a frozen VLA at inference time without failure data, is proposed, a training-free framework that recovers a frozen VLA at inference time without policy fine-tuning or failure-specific recovery training.

Yan-Yan Zhang, Di-Sheng Liu, Kai Ye et al. · 2 citations
Preprint Aug 2026

WA-SpecDec: World-Aware Speculative Decoding for Vision-Language-Action Models

WA-SpecDec is proposed, a world-aware speculative decoding framework that injects world-model-derived physical scene awareness during the VLA prefill stage, producing shared world-aware prefill states for draft proposal and target verification without changing the relaxed acceptance rule.

Zikang Wen, Yuning Zhang, Dong Yuan · 0 citations
Preprint Sep 2026

SafeStage: Evaluating Safety Before, During, and After Vision-Language-Conditioned Robot Manipulation

Vision-language-conditioned robot policies integrate perception, language understanding, and control for general-purpose manipulation. However, existing evaluations often focus on task success, isolated physical constraints, semantic refusal, or realized physical damage, providing limited insight into where safety fail...

Jin-Zhu Luo, Qi Zhang, Wei Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.