
YZ
Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana
· 1 min read
ResearcharXiv cs.AI
Optimal Transport Meets Reinforcement Learning: A Survey
arXiv:2610.01413v1 Announce Type: cross
Abstract: Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, commonly used divergences may become ineffective when these distributions overlap weakly, which is frequently encountered in imitation learning, offline RL, and deployment under distribution shift. Optimal transport (OT) offers an alternative by measuring the cost of \emph{moving} probability mass from one distribution to another under a ground cost that encodes task geometry. This survey covers how OT is used inside RL objectives and algorithms. For each method, we identify: the role OT plays, the distributions compared, the OT formulation used, and the treatment of temporal structure. Beyond categorising existing methods, we discuss the motivations behind different OT choices, practical considerations such as cost design and computational challenges, and highlight open problems including scalable trajectory-level transport, principled handling of mass mismatch, and theoretical analysis for OT-regularised RL.
Original source
This story was published by arXiv cs.AI and written by Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


