
MZ
Mingyuan Zhang, Yue Bai, Yifan Wang, Yiyang Huang, Yun Fu
· 1 min read
ResearcharXiv cs.CV
Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
arXiv:2512.23073v2 Announce Type: replace-cross
Abstract: Fine-tuning has become the dominant paradigm for adapting Vision-Language Models (VLMs), yet most approaches rely on explicit weight updates that introduce a fundamental trade-off. Full Fine-Tuning (FFT) may perturb pretrained representations due to cross-modal gradient interference, whereas Parameter-Efficient Fine-Tuning (PEFT) methods rely on additive modules, such as low-rank adapters, which may limit adaptation capacity. In this paper, we rethink VLM adaptation from a structural selection framework that adapts VLMs without modifying backbone weights, and we propose Mask Fine-Tuning (MFT). MFT learns masks that selectively route information through existing pretrained connections, dynamically uncovering subnetworks that better align pretrained representations with downstream objectives. Extensive experiments show that MFT provides an effective structural alternative to both FFT and PEFT, consistently achieving superior performance across multiple vision-language benchmarks without adding knowledge or altering the deployment architecture. Moreover, our analysis with MFT provides new insights into how pretrained VLMs reorganize their internal representational pathways during adaptation.
Original source
This story was published by arXiv cs.CV and written by Mingyuan Zhang, Yue Bai, Yifan Wang, Yiyang Huang, Yun Fu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


