Skip to content
Preprint

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies

Sep 2026 · 0 citations · 55 references
Computer Science

TL;DR

This work proposes CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution, and introduces the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anomalies.

Abstract

Vision-Language-Action (VLA) policies achieve strong performance in robotic manipulation but remain brittle once execution deviates from nominal trajectories. We propose CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution. Instead of generating corrective data from manually designed or random perturbations, CARE collects failed rollouts, models stage-conditioned post-failure deviations, and uses the resulting empirical distributions to synthesize representative failure states and corrective demonstrations. At inference time, CARE combines stage-wise planning with physically grounded 3D monitoring to trigger atomic adjustments or re-operations while preserving task progress. We further introduce the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anomalies. Experiments across multiple VLA backbones, simulation benchmarks, and real-world dual-arm tasks show consistent improvements, with average task-success gains of 14.5 points in simulation and 15.9 points in the real world. Code, models, and data are available at https://github.com/xiaojunlan/care

View source

Similar papers

Preprint Sep 2026

VLA-Corrector: Stage-Aware Observable State Understanding for Prompt-Based Closed-Loop Recovery of Vision-Language-Action Policies

Results show that observable visual-proprioceptive-action history is sufficient to infer latent task states and enable practical failure recovery for existing VLA policies, and the proposed approach substantially improves closed-loop reliability under diverse perturbations.

Chang-Nan Song, Bin Qian, Yan Feng et al. · 0 citations
Preprint Oct 2026

Recova: Agent-Guided Failure Recovery for Autonomous Robotic Manipulation

Manipulation failures can leave scenes in states from which a task policy cannot recover. Learning corrective behaviors requires scalable failure exploration and physical grounding. We present Recova, an agent-guided framework that jointly develops task execution and recovery in a reconstructed digital twin, then verif...

Isabella Liu, An-Chieh Cheng, Johan Bjorck et al. · 0 citations
Preprint Aug 2026

Imagining Recovery: Inference-Time Counterfactual Realignment for Vision-Language-Action Models

Counterfactual Realignment (CoRe), a training-free framework that recovers a frozen VLA at inference time without failure data, is proposed, a training-free framework that recovers a frozen VLA at inference time without policy fine-tuning or failure-specific recovery training.

Yan-Yan Zhang, Di-Sheng Liu, Kai Ye et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent pol...

M.-Y. Cui, Zhe-Yuan Liu, Yi-Han Zhu et al. · 0 citations
Preprint Sep 2026

ProAct-VLM: Pre-Failure Vision-Language Task Replanning with Continuous Perception Feedback

Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly. Traditional rule-based decision-making pipelines are brittle...

Ahmed N. Ahmed, Omar Moured, Mughni Irfan Mohammed Abdul et al. · 0 citations
Preprint Sep 2026

Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

FIND is introduced, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace and reframes autonomous practice as a scene-conditioned, performance-aware task-selection problem.

Yuan Fang, Ze-Chu Li, Hao-Lei Tong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.