Skip to content
Open access

Predictive processing as a scalable computational principle for embodied multitask intelligence

Aug 2026 · Science Advances · Vol 12 · 1 citation · 54 references
Medicine

TL;DR

A scalable hierarchical multimodal recurrent neural network grounded in predictive processing under the free-energy principle, capable of directly integrating more than 30,000-dimensional visuo-proprioceptive inputs without dimensionality reduction or handcrafted preprocessing is introduced.

Abstract

Humans exhibit remarkable flexibility in adapting to diverse and uncertain environments—a hallmark arising from the brain’s ability to integrate multimodal sensory streams into coherent predictive models. Drawing on this principle, we introduce a scalable hierarchical multimodal recurrent neural network grounded in predictive processing under the free-energy principle, capable of directly integrating more than 30,000-dimensional visuo-proprioceptive inputs without dimensionality reduction or handcrafted preprocessing. Using sensory data from teleoperation of a full-scale physical humanoid robot performing two caregiving-related tasks—rigid-body repositioning and flexible-towel wiping—the model learns to predict high-dimensional visuo-proprioceptive streams end to end. In open-loop adaptive inference experiments, the framework exhibits three emergent properties: (i) self-organized hierarchical latent dynamics governing task transitions, uncertainty, and occlusion inference; (ii) robustness to degraded vision via multimodal integration; and (iii) asymmetric interference in multitask learning. Although evaluated in simulations, the framework is extensible to closed-loop robot control, with proprioceptive predictions driving action, thereby establishing a generalizable computational foundation bridging brain theory, artificial intelligence, and embodied robotics.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Brain-Inspired Hierarchical Modularity for General Continual Learning

Across visual recognition, vision-language understanding, ego-exo video understanding, and embodied vision-language-action learning, this method consistently improves learning under online and uncertain data streams, with gains exceeding 50 percentage points over replay-free alternatives in embodied manipulation.

Hong-Wei Yan, Kang-Lei Zhou, Qi-Hao Cheng et al. · 0 citations
Open access Sep 2026

Human-like memory empowers embodied robots for long-term object navigation

Effortless object finding by humans, even in cluttered or unseen environments, relies on the seamless integration of perception, memory, and contextual inference. In contrast, embodied robots operating under egocentric perception and partial observability frequently struggle with dynamic spatial relations and long-te...

Ying Zhang, Ren-Jie Song, Hong-Liang Ren et al. · 1 citation
Preprint Sep 2026

Generalist Open-World Temporal Perception

The next generation of artificial intelligence systems will likely be natively temporal and multimodal in both inputs and outputs: able to converse, perceive, predict, reason, and synthesize through a shared world representation. Realizing this requires a temporal perceptual substrate integrating sensory streams, langu...

C. Sminchisescu · 0 citations
Preprint Sep 2026

ME-Brain-1.0: Memory, Cognition and Action for Evolving Embodied Intelligence

MachEmbodied-Brain (ME-Brain) is introduced, a self-evolving embodied system organized around a closed loop of action execution, experience acquisition, experience evolution, and improved execution that shifts embodied intelligence from train-and-freeze to deploy-and-evolve without model retraining.

Wei He, Heng-Tao Li, Chen-Feng Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Learning Options for Compositional Motor Control with Adapter Banks

Learning flexible motor primitives is a hallmark of skilled motor control. Recent neuroscience theory proposes that motor primitives may be implemented as low-rank perturbations of a shared recurrent network, but leaves open how such a system is learned. We translate this principle into a novel architecture for learnin...

Sreejan Kumar, M. Mattar, Lea Duncker · 0 citations
Preprint Aug 2026

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantial...

GigaBrain Team, An-Gen Ye, Axiang Sun et al. · 5 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.