
PR
Panagiotis Roditis, Panagiotis P. Filntisis, Petros Maragos
· 1 min read
ResearcharXiv cs.AI
A Geometric Approach to Soft Actor-Critic with Zonotopes for Locomotion Learning
arXiv:2610.12113v1 Announce Type: cross
Abstract: Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics. This uses the same aggregation rule everywhere, regardless of how the critics disagree. We propose \textbf{GeZo-SAC}, which uses auxiliary geometric representations to adapt critic pessimism to the state and action. Alongside its scalar value, each critic predicts a set of generators defining a zonotope. Probing this zonotope along sampled directions provides a geometric width, "subtracted from each critic value as a pessimistic offset, and a measure of disagreement between the two critics, aggregated with log-sum-exp. This disagreement controls how the critics are combined, moving from a width-weighted average toward the usual minimum as disagreement increases. At inference, the deployed policy is an unmodified SAC actor, since the generators are used only on the critic side during training.Across four MuJoCo-v5 locomotion benchmarks and six off-policy baselines, GeZo-SAC achieves the highest mean return on Ant-v5 and Hopper-v5 and remains competitive with other methods on the remaining tasks. Our analysis further shows that GeZo-SAC achieves the lowest average actuator work and action effort per metre among the evaluated methods, while maintaining near-zero measured overestimation frequency across all four environments.
Original source
This story was published by arXiv cs.AI and written by Panagiotis Roditis, Panagiotis P. Filntisis, Petros Maragos. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


