SyncAI.news, a Varaisys broadcasting
Uncertainty-Normalized Margins for Direct Preference Optimization
SK

Sadegh Khorasani, Petrus Mikkola, Matthias Grossglauser

· 1 min read

ResearcharXiv cs.LG

Uncertainty-Normalized Margins for Direct Preference Optimization

arXiv:2609.38647v1 Announce Type: new Abstract: Direct preference optimization (DPO) models binary preferences through a Bradley-Terry model with a common noise scale, without explicitly accounting for preference strength or prompt-dependent uncertainty from human feedback. We introduce uncertainty-normalized margin DPO (UNM-DPO), which combines strength-dependent margins with a learned prompt scale. Motivated by a heteroskedastic Bradley-Terry model, we develop two training objectives. Both compare the implicit rewards of preferred and rejected responses, derived from response log-probability ratios to a reference policy. Advantage-only (AO) divides this reward difference by the prompt scale before subtracting the margin; whole-residual (WR) subtracts the margin before dividing by the scale. For the WR comparison model, we establish a necessary and sufficient condition under which known margins make the prompt scale identifiable. We introduce a practical procedure for learning the scale. Building on WR, we introduce ULNM-DPO-WR, which normalizes each response's implicit reward by its length. We evaluate our methods against DPO and related baselines on HelpSteer2 and HelpSteer3, using the Skywork reward model as a judge. With Llama-3.1-8B-Instruct, ULNM-DPO-WR achieves tie-adjusted win rates against matched DPO of 68.00% and 65.31% on evaluation panels, with higher mean rewards and shorter responses on average. On AlpacaEval with a GPT-4.1 judge and GPT-4-Turbo reference answers, the same 8B policy achieves a length-controlled win rate of 21.62%, compared with 16.39% for DPO and 15.30% for SimPO. These results demonstrate the potential of combining preference-strength margins, learned prompt scales, and length normalization for policy optimization.

Original source

This story was published by arXiv cs.LG and written by Sadegh Khorasani, Petrus Mikkola, Matthias Grossglauser. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News