Skip to content

A computational framework for Kármán gaiting in robotic fish: spatio-temporal perception and CPG-based reinforcement learning

Aug 2026 · Bioinspiration & Biomimetics · Vol 21 · 1 citation · 56 references
Medicine Physics

TL;DR

A fully computational framework focusing on the modeling and simulation of a spatio-temporal sensory system to autonomously generate the Kármán gait is proposed, providing a robust algorithmic blueprint for future physical deployments in complex aquatic environments.

Abstract

Navigating in unsteady wake flows, such as Kármán vortex streets, presents a formidable challenge for biomimetic autonomous underwater vehicles. Biological fish achieve this by utilizing their lateral line sensory systems to perceive local flow gradients and adopting an energy-efficient swimming pattern known as the Kármán gait. To translate this biological phenomenon into a practical robotics engineering solution, this paper proposes a fully computational framework focusing on the modeling and simulation of a spatio-temporal sensory system to autonomously generate the Kármán gait. To overcome the unrealistic assumption of full-state observability common in existing reinforcement learning studies, we model a multi-point lateral line array coupled with a frame-stacking mechanism. This allows the simulated agent to reconstruct the spatio-temporal topology of the surrounding unsteady flow relying exclusively on local pressure and velocity gradients. The sensory model is integrated with a spatio-temporal perceptual twin delayed deep deterministic policy gradient (STP-TD3) algorithm, which drives a Hopf-oscillator-based central pattern generator. Through rigorous high-fidelity computational fluid dynamics simulations, we quantitatively evaluate the autonomous emergence of the Kármán gait by assessing the agent’s kinematic energy proxy-mapped from joint actuation effort. Results reveal that the agent expends significantly less mechanical effort navigating through the turbulent vortex street compared to swimming in steady water, suggesting the active exploitation of the local wake dynamics. The results theoretically underscore the necessity of distributed STP for biomimetic robots, providing a robust algorithmic blueprint for future physical deployments in complex aquatic environments.

View source

Similar papers

Open access Aug 2026

Adaptive Task-Oriented Locomotion Control of a 2D Planar Robotic Fish Model Using Deep Reinforcement Learning and Sensory-Feedback CPG Network

Autonomous locomotion in robotic fish requires task-dependent control capabilities under changing environmental conditions. This paper proposes a hierarchical simulation-based control framework for a two-joint robotic fish in a two-dimensional (2D) planar environment. This framework integrates the twin delayed deep det...

Gonca Ozmen Koca, D. Korkmaz, Cafer Bal et al. · 0 citations

Adaptive Undulatory Locomotion of Snake-like Robots in Dynamic Viscous Environments via Deep Reinforcement Learning

This paper demonstrates how deep reinforcement learning (DRL) enables adaptive locomotion of snake-like robots in dynamically changing viscous environments, overcoming the inherent performance limitations of classical predefined control methods. The lack of direct onboard sensors for fluid properties necessitates formu...

T. Kimoto, A. Yamano, Kohei Honda et al. · 0 citations
Jul 2026

Egocentric Station Holding of Robotic Fish in Unknown Turbulent Background Flow

Approaching a target position and holding station in flowing water is a fundamental and critical capability for robotic fish operating in natural aquatic environments. Despite decades of advances in enhancing swimming efficiency and maneuverability, this capability remains underdeveloped, largely owing to the insuffici...

Xiaozhu Lin, Xuejiao Huang, Hongru Dai et al. · 0 citations
Conference Aug 2026

End-to-End Control of a Quadruped Robot Using Deep Reinforcement Learning

This paper presents the development and implementation of an end-to-end control framework for a quadruped walking robot based on deep reinforcement learning. The primary objective of the study is to design and verify a control system capable of autonomously generating locomotion strategies. A model of the walking robot...

Filip Połatyński, Paweł Skruch · 0 citations
Jul 2026

Reinforcement Learning on Cost-Constrained Quadrupedal Hardware

Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-to-real gap. The chasm of simulation to deployment in hardware lies in the delay of the actuator reaching the commanded position. On platforms such as the Mini Pupp...

Javier C. Weddington, Bence P. Ölveczky, S. Baccus · 0 citations
Preprint Aug 2026

Spatiotemporal Agility: Time-Constrained Reinforcement Learning for Vision-Guided Dynamic Quadrupedal Interception

An integrated framework that combines a vision module for landing point and time prediction with a direct position and time conditioned RL locomotion policy, instead of intermediate velocity commands is proposed, which mitigates perception latency during dynamic interception.

Yi-Dong Zhu, Zibo Dai, Tong-Ning Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.