
SY
Sin-Han Yang, Shih-Cheng Huang, Chieh-Yen Lin, Yun-Nung Chen, Shao-Hua Sun, Hung-yi Lee
· 1 min read
ResearcharXiv cs.CL
A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic
arXiv:2610.07990v1 Announce Type: cross
Abstract: Model merging aims to build a multi-task model cheaply by combining the weights of individual task-specific models. To perform well across multiple tasks, most existing merging methods use an additional dataset to find the coefficients for the best linear combination of task-specific weight updates. However, we identify an implicit regularization in this standard practice: searching over coefficients restricts the candidate models to a subspace spanned by task-specific weight updates. In this work, we investigate whether this regularization is actually useful. Surprisingly, empirical results show that optimizing merged-model weights without this regularization significantly boosts the performance of common merging methods across multiple architectures, domains, and even in an extremely data-limited scenario where only one instance is available per class. Moreover, directly optimizing the pretrained model weights even outperforms some existing merging methods. Analysis shows that better multi-task weights exist outside the subspace and can be found using multiple methods. We study different strategies for using the additional dataset, discussing their practical use and implications for model merging. Overall, this work calls for revisiting the existing model-merging pipeline, motivating a broader exploration of the weight space and a reconsideration of the implicit regularization induced by task arithmetic.
Original source
This story was published by arXiv cs.CL and written by Sin-Han Yang, Shih-Cheng Huang, Chieh-Yen Lin, Yun-Nung Chen, Shao-Hua Sun, Hung-yi Lee. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


