Variational Model Merging for Pareto Front Estimation in Multitask Finetuning

작성자

카테고리:

← 피드로
arXiv cs.AI · Hugo Monz'on Maldonado, Nico Daheim, Thomas M"ollenhoff, Iryna Gurevych, Mohammad Emtiyaz Khan · 2026-06-24 AI

[Submitted on 11 Dec 2024 (v1), last revised 23 Jun 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Pareto fronts are useful to find good task-mixing strategies for multitask finetuning, but they are also costly to compute. To reduce costs, recent works have used existing model merging methods to help train cheap surrogate models to estimate the Pareto fronts. However, no work has yet considered designing new model-merging methods to directly, and provably, improve the quality of Pareto fronts. Here, we fill this gap by proposing a new Bayesian approach called Variational Model Merging. In this approach, existing model-merging methods are obtained as special cases of “posterior-merging” when Gaussian posteriors are used and new model-merging strategies can be derived by using non-Gaussian posteriors. Our main theoretical result is to show that more flexible posteriors necessarily yield better estimates of Pareto fronts. For instance, a Pareto front estimate obtained by merging full-Gaussian posteriors is expected to be better than that obtained by using isotropic Gaussian posteriors. We validate the theory through extensive empirical results on vision and language transformers where better Gaussian families consistently yields better or comparable Pareto fronts. Our work is a rare instance where Bayesian ideas are used to improve Pareto analysis.

Submission history

From: Hugo Monzón Maldonado [view email]
[v1] Wed, 11 Dec 2024 07:06:36 UTC (1,232 KB)
[v2] Tue, 23 Jun 2026 16:19:45 UTC (2,662 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2412.08147

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다