SyncAI.news, a Varaisys broadcasting
DOHF: Online Diffusion Fine-tuning with Doob's $h$-transform Guidance
ZG

Zhengyi Guo, Jiayuan Sheng, Wenpin Tang

· 1 min read

ResearcharXiv cs.LG

DOHF: Online Diffusion Fine-tuning with Doob's $h$-transform Guidance

arXiv:2609.31882v1 Announce Type: new Abstract: Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to estimate. In this work, we propose Diffusion Online $h$-guidance Fine-tuning (DOHF), which turns Doob's $h$-transform into a practical online training algorithm. DOHF assigns optimality weights to generated samples, estimates the normalized local correction $\nabla\log h$ under the current rollout policy, and distills it directly into the generative model. Theoretically, we characterize the population-optimal DiffusionNFT update as well as the various classfier free guidance methods through a unified $h$-transform perspective. Methodologically, our framework accommodates black-box and non-differentiable rewards without additional network evaluations. We further show improved alignments under three empirical scenarios. Our work demonstrates how adapting probabilistic conditioning through inexpensive estimation and iterative distillation can improve generative learning across statistical sampling and visual generation.

Original source

This story was published by arXiv cs.LG and written by Zhengyi Guo, Jiayuan Sheng, Wenpin Tang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News