Training-free model merging combines multiple task-specialized models into a single multi-task
model without joint retraining. Existing methods merge weight matrices $\mathbf{W}_Q, \mathbf{W}_K,
\mathbf{W}_V, \mathbf{W}_O$ independently — ignoring the fundamental structure of
transformer attention, where behavior depends on the composed operators
$\mathbf{W}_Q\mathbf{W}_K^\top$ (which inputs the model attends to) and
$\mathbf{W}_V\mathbf{W}_O$ (what information is written to the residual stream).
We propose Sequential Local Operator Alignment (SLOA), which merges these
local functional operators directly. Each operator is aligned via a closed-form
regularized least-squares objective. Crucially, alignment is performed
sequentially in topological order — each step uses activations from the
partially-merged model, correcting for accumulated upstream errors. After merging in the
ambient $d \times d$ space, each operator is factorized back via SVD at a chosen rank $r$,
allowing rank expansion ($r > d_k$) as a capacity-control mechanism that
retains multi-task singular directions without any training.
SLOA achieves state-of-the-art training-free merging on CLIP/ViT vision benchmarks (8, 14,
and 20 tasks) and RoBERTa NLU benchmarks (7 GLUE tasks), outperforming RegMean++ by
+6.6 pp on the hardest 20-task vision setting.