SyncAI.news, a Varaisys broadcasting
Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning
TL

Toyota Li, David Zhao, Alan Zhao

· 1 min read

ResearcharXiv cs.LG

Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning

arXiv:2609.33444v1 Announce Type: new Abstract: A nascent family of methods that forgoes the policy gradient and reweights a supervised regression instead has garnered momentum in reinforcement learning for diffusion and flow models. DiffusionNFT, FlowAWR, and RAM are representative regimes with contrasting motivations. It is yet opaque what, if anything, they share. We substantiate that each is the solution of one divergence-constrained reward-maximization problem, and they are differentiated only by the convex generator that defines the constraint. Under the unified modeling framework, we unravel the relaxations that prior art made during building the advantage-embedded regression target: approximating the KKT condition and posterior normalizer for the linear and exponential tilt shapes DiffusionNFT and FlowAWR respectively, while preserving the exact sparsemax projection onto the probability simplex for linear tilt leads to another superior model type in this work. Beyond the theoretical underpinnings, we further empirically investigate the design space and shed light on the training recipe for regression-style diffusion RL. Retaining the merits discovered during our exploration gives rise to DiffusionRFT, our paradigm that converges faster, trains more stably, and attains the top performance.

Original source

This story was published by arXiv cs.LG and written by Toyota Li, David Zhao, Alan Zhao. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News