Skip to content

Adaptive Undulatory Locomotion of Snake-like Robots in Dynamic Viscous Environments via Deep Reinforcement Learning

Jul 2026 · arXiv.org · Vol abs/2607.21960 · 0 citations · 27 references
Computer Science

Abstract

This paper demonstrates how deep reinforcement learning (DRL) enables adaptive locomotion of snake-like robots in dynamically changing viscous environments, overcoming the inherent performance limitations of classical predefined control methods. The lack of direct onboard sensors for fluid properties necessitates formulating this task as a partially observable Markov decision process. By employing an asymmetric actor-critic framework, a teacher policy trained using privileged information available only in the physics simulator distills its knowledge into a student policy that relies solely on proprioceptive sensor information. Simulation results across a wide range of dynamic viscosity changes ($10^{-7}$ to $10^{-2} m^2/s$) reveal that the DRL agent autonomously acquires non-sinusoidal adaptive gaits. These gaits improve propulsion velocity and transport efficiency, breaking the inherent limits of conventional sinusoidal and kinematic control. The findings establish that implicit environment inference via privileged information distillation is an effective approach to bypass the constraints of classical models under unpredictable fluid dynamics.

View source

Similar papers

Open access Jul 2026

From insect behavior to transferable robot locomotion: inferring embodied locomotor principles from limited data via adversarial inverse reinforcement learning

Insect locomotion exhibits remarkable adaptability and flexibility despite the limited scale of its nervous system. However, the underlying principles that govern leg coordination remain difficult to extract and model computationally. Understanding how insects achieve stable and adaptive locomotion has long provided im...

Yuchen Wang, Thirawat Chuthong, M. Hayashibe et al. · 0 citations
Review Open access Sep 2026

Extending the Speed Limit of Quadrupedal Locomotion via Refined Actuator Modeling and Adaptive Command Scheduling

A refined actuator model explicitly captures high-speed voltage coupling and magnetic saturation, enabling a more accurate representation of the torque–speed envelope and a reinforcement learning framework incorporating a two-stage curriculum and adaptive command scheduling (ACS) ensures stable training.

Yu-Cheng Tao, Shao-Wen Cheng, Guo-Rong Lan et al. · 0 citations
Aug 2026

A computational framework for Kármán gaiting in robotic fish: spatio-temporal perception and CPG-based reinforcement learning

A fully computational framework focusing on the modeling and simulation of a spatio-temporal sensory system to autonomously generate the Kármán gait is proposed, providing a robust algorithmic blueprint for future physical deployments in complex aquatic environments.

Xin-Qi Wang, Ming Wang, Xin-Yan Liu et al. · 1 citation
Conference Aug 2026

End-to-End Control of a Quadruped Robot Using Deep Reinforcement Learning

This paper presents the development and implementation of an end-to-end control framework for a quadruped walking robot based on deep reinforcement learning. The primary objective of the study is to design and verify a control system capable of autonomously generating locomotion strategies. A model of the walking robot...

Filip Połatyński, Paweł Skruch · 0 citations
Open access Aug 2026

Application and evaluation of reinforcement learning for two-dimensional trajectory tracking in snake-like robots

Context—Snake-like robots are biomimetic systems that can move effectively in narrow, complex, and restricted environments thanks to their modular and flexible body structures composed of numerous serially connected joints. These characteristics offer significant advantages, particularly in areas such as pipeline inspe...

Furkan Mezgil, M. Bingöl · 0 citations
Preprint Aug 2026

Learning Fault-Tolerant Locomotion with Adaptive Gait Timing

A deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss that employs an asymmetric actor-critic architecture in which the critic has access to privileged information during training, while the actor learns to reconstruct a corresponding latent representation from proprioceptive ob...

Giovanbattista Gravina, Luca Rossini, Carlo Rizzardo et al. · 1 citation

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.