Conceptual

Multi-Task Model Merging via Adaptive Projective Gradient Descent

A data-free method that merges independently fine-tuned expert models into one multi-task model by framing merging as a constrained optimization problem: minimize the gap between the merged model and each task-specific model while preserving knowledge shared across tasks. It projects task-vector gradient updates into a subspace shared by all tasks and treats the per-task merging coefficients as adaptive, training-free learning rates.