
DY
Daisuke Yamada, Travis Pence, Vikas Singh
· 1 min read
ResearcharXiv cs.LG
Learning Goal-Reaching Quasimetric Geometry From Finite-Time Reachability
arXiv:2610.00778v1 Announce Type: new
Abstract: In goal-conditioned reinforcement learning (GCRL), quasimetric learning models goal-reaching costs as quasimetric distances, connecting local constraints to global value geometry. Its local constraints, however, should reflect the direction- dependent effects of control composition over a finite horizon together with environmental feasibility. We propose ReQRL, which constrains the critic's value gradients through finite-horizon reachability. Drawing on state-constrained optimal control, we decouple dynamical reachability from boundary geometry, estimating both from data. On OGBench, our method outperforms or rivals existing quasimetric approaches and other offline GCRL methods.
Original source
This story was published by arXiv cs.LG and written by Daisuke Yamada, Travis Pence, Vikas Singh. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


