
ZY
Zikun Ye, Jinglong Zhao, Lei Wang
· 1 min read
ResearcharXiv cs.LG
Fine-Tune, Then Rectify
arXiv:2511.19486v3 Announce Type: replace
Abstract: Driven by recent advances in artificial intelligence, a growing literature has demonstrated the potential of using large language models (LLMs) as scalable surrogates to generate human-like responses. Two common approaches to improve the performance of LLMs include: fine-tuning, which aligns the LLM more closely with human responses, and rectification, which corrects biases in LLM outputs. In this paper, we develop a two-stage framework that combines fine-tuning and rectification, and optimally allocates limited labeled samples across the two stages. A key insight is that the conventional fine-tuning objective of minimizing mean squared prediction error is generally not aligned with the downstream rectification stage. For mean estimation, we propose to minimize the variance of the prediction errors; for general M-estimation, we propose to minimize a scalarized variance metric as the fine-tuning objective. Building on this insight, we leverage the scaling law of fine-tuning to optimally allocate the limited labeled human data between the fine-tuning and rectification stages. Our empirical analysis validates the fine-tuning scaling law and confirms that our proposed optimal allocation rule reliably identifies the optimal sample allocation. We demonstrate substantial efficiency gains in estimation and inference performance relative to fine-tuning or rectification alone, or to employing the conventional mean squared error objective within the fine-tuning then rectification framework. Such efficiency gains translate to significant cost savings for making reliable decisions.
Original source
This story was published by arXiv cs.LG and written by Zikun Ye, Jinglong Zhao, Lei Wang. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


