SyncAI.news, a Varaisys broadcasting
Hierarchical Reinforcement Learning for Collision-Free Locomotion of an Underactuated Biped
JP

Jagannath Prasad Sahoo, Saurabh Kumar, Surya Prakash S. K., Samiran Datta, Abhay Dwivedi, Amit Shukla

· 1 min read

ResearcharXiv cs.AI

Hierarchical Reinforcement Learning for Collision-Free Locomotion of an Underactuated Biped

arXiv:2610.05855v1 Announce Type: cross Abstract: A bipedal robot cannot deviate from its path to avoid an obstacle without disturbing its balance, and this coupling is most severe on underactuated platforms such as the biped considered here, which has four actuated joints per leg and no hip or ankle roll. This paper presents a Hierarchical Reinforcement Learning (HRL) framework in which a High-Level (HL) policy observes the robot pose, 36 raycast proximity measurements, moving-obstacle states, and a receding-horizon local goal, and outputs a body-velocity command $(v_x, v_y, \omega_{yaw})$ every ten control steps, while a velocity-conditioned Low-Level (LL) policy tracks each command through PD-controlled joint targets. Both policies are trained jointly with Soft Actor-Critic (SAC). Because the converged gait is task-agnostic, it is frozen and driven by classical planners over the same command interface, yielding three controlled baselines: SAC+A*, SAC+RRT*, and SAC+APF. Across 100 evaluation trials per method in randomized PyBullet environments, the proposed method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the planner hybrids, with path lengths within 4% of the A* reference, and ablations confirm that each observation channel and reward term contributes materially to this performance.

Original source

This story was published by arXiv cs.AI and written by Jagannath Prasad Sahoo, Saurabh Kumar, Surya Prakash S. K., Samiran Datta, Abhay Dwivedi, Amit Shukla. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News