Skip to content

WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving

Jul 2026 · arXiv.org · Vol abs/2607.08375 · 1 citation · 65 references
Computer Science

TL;DR

WCog-VLA is proposed, a novel dual-level World-Cognitive VLA framework that successfully bridges semantic world forecasting with generative world evolution to achieve proactive autonomous driving.

Abstract

Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving. However, existing methods either lack comprehensive world cognition or suffer from fragmented world foresight, inherently confining these models to reactive driving. To address this limitation, we propose WCog-VLA, a novel dual-level World-Cognitive VLA framework that successfully bridges semantic world forecasting with generative world evolution to achieve proactive autonomous driving. At the semantic level, WCog-VLA unifies world cognition and reasoning by incorporating 3D spatial perception and injecting agent tokens to capture the world dynamics, while concurrently enabling Game-theoretic Chain-of-Thought (Game-CoT) reasoning. At the generative level, we introduce the Aligned Decoupled Diffusion Transformer (ADDT) as a powerful generative world model that synthesizes physically-plausible joint multi-agent trajectories. Through scene representation alignment, ADDT reduces the number of denoising steps required and thus significantly accelerates inference. To facilitate strategic reasoning, we further construct a large-scale dataset featuring 85k Game-CoT annotations. Extensive experiments on the NAVSIM benchmark demonstrate that WCog-VLA achieves a State-Of-The-Art (SOTA) PDMS score of 92.9.

View source

Similar papers

Aug 2026

Towards a Causally-inspired Evolving World Model for Vision-and-Language Navigation in Continuous Environments.

Vision-and-Language Navigation in Continuous Environments (VLN-CE) has emerged as a pivotal challenge in Embodied AI, requiring an agent to navigate 3D spaces guided by natural language instructions. Drawing inspiration from human cognition, world models provide a powerful paradigm by predicting environment dynamics and enabling reasoning beyond immediate observations. However, existing world model-based VLN methods remain static once trained - their representations rely on fixed correlationbased priors rather than adaptive causal structures, making them unable to accommodate evolving confounders and changing observation-action dependencies across environments. This rigidity leads to overfitting to training-specific patterns and degraded performance under distribution shifts. To address this limitation, we propose a causally-inspired evolving world modeling framework that formulates VLN-CE as a sequence of causal partially observable Markov decision processes. Our model learns unified latent states that integrate vision, language, and action, while addressing spurious correlations through a dual-level intervention mechanism: at the observation level, frequency-domain perturbations simulate superficial appearance variations to enhance perceptual robustness; at the representation level, cross-episode confounder buffers perform counterfactual substitution to approximate the influence of latent confounding factors. Beyond static world modeling, our framework continuously evolves, refining these proxy representations across episodes, enabling efficient adaptation to previously unseen environments. Building on this evolving causally-inspired foundation, our world model supports counterfactual reasoning and strengthens generalization across diverse navigation contexts. Extensive evaluations on established VLN-CE benchmarks demonstrate that our method outperforms existing approaches, delivering superior navigation performance across diverse scenarios. Real-world robot evaluations further validate the practicality of our approach. Code is available in the Supplementary Material.

Xuan Yao, Junyu Gao, Changsheng Xu · 0 citations
Preprint Aug 2026

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.

Guiyu Zhao, Longteng Guo, Yanghong Mei et al. · 0 citations
#artificial intelligence Preprint Aug 2026

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

This work proposes BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations, and introduces an asynchronous rectified-flow inference strategy with decoupled video and action denoising.

Bing Zhan, Shuyao Shang, Shuo Lu et al. · 1 citation
Preprint Aug 2026

XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving

This work proposes XCoT-VLA, which replaces descriptive rationales with compact executable CoT tokens learned from automatically constructed Reason-Action supervision, and demonstrates that driving-oriented reasoning can be compact, executable, and directly connected to trajectory generation.

Foundation Model Team, XPeng Inc · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.