Robots operating in open environments act under partial observability, physical constraints, and dynamic task contexts. Beyond mapping observations and language instructions to actions, they must anticipate how candidate actions may affect future states and task-relevant outcomes. Recent advances in world models, video generation, and Vision-Language-Action (VLA) policies have motivated the development of World-Action Models (WAMs), which couple future world prediction with executable action generation. This survey provides a robotics-oriented review of WAMs. We clarify their scope relative to conventional world models, model-based reinforcement learning, action-conditioned video generation, and reactive VLA policies, and organize existing methods through a unified taxonomy covering representations, transition modeling, action interfaces, architectures, training pipelines, data modalities, and scaling strategies. We further review applications of WAMs in manipulation, navigation, and autonomous driving, and we summarize the datasets, benchmarks, metrics, and protocols used to evaluate WAM systems. Finally, we discuss key challenges in action alignment, world-action factorization, spatial and multi-view consistency, long-horizon memory, neural simulation for closed-loop policy learning, and efficient inference. Taken together, this survey aims to provide a concise technical foundation for integrating predictive world modeling with action generation, toward more reliable embodied robot intelligence. Project page: https://rcl-robotics.github.io/Awesome-World-Action-Models.
Zu-Xing Lu, Hong-Jia Zhai, Guan-Zhi Wang et al.· 0 citations
Electronic skin powered with artificial intelligence could enable next-generation robotic and medical devices, yet integrating multimodal sensors and analyzing heterogeneous, multifrequency time series remain challenging. Most wearable machine learning architectures are time-invariant and trained for a specific task, limiting transfer across modalities and users. We present a multimodal electronic skin that captures diverse physiological signs with an adaptive learning framework that rapidly generalizes to unseen tasks with minimal labeled data. Our streamlined end-to-end framework uses a spectral variational autoencoder to denoise and compress multifrequency biosignals into a shared, unified second-wise latent space that preserves the spectral-temporal structure, followed by a transformer to capture temporal dependencies to support diverse downstream tasks with data-efficient learning. We demonstrate robust adaptation with 94.7% accuracy in activity recognition and 90.2% precision in fatigue assessment across various users and daily activities regardless of device and user variations, highlighting a scalable route to generalized physiological time-series analytics and human performance assessments.