Sequential Local Operator Alignment
for Training-Free Model Merging

1CISPA Helmholtz Center for Information Security   1Universität des Saarlandes   2Technical University of Munich   3Munich Center for Machine Learning   4Massachusetts Institute of Technology
SLOA method illustration

SLOA aligns the effective local operators ($\mathbf{W}_Q\mathbf{W}_K^\top$, $\mathbf{W}_V\mathbf{W}_O$, LayerNorm affine, MLP projections) rather than raw weight matrices, sequentially under activations of the partially-merged model, then factorizes each operator back into transformer-compatible weights.

Abstract

Training-free model merging combines multiple task-specialized models into a single multi-task model without joint retraining. Existing methods merge weight matrices $\mathbf{W}_Q, \mathbf{W}_K, \mathbf{W}_V, \mathbf{W}_O$ independently — ignoring the fundamental structure of transformer attention, where behavior depends on the composed operators $\mathbf{W}_Q\mathbf{W}_K^\top$ (which inputs the model attends to) and $\mathbf{W}_V\mathbf{W}_O$ (what information is written to the residual stream).

We propose Sequential Local Operator Alignment (SLOA), which merges these local functional operators directly. Each operator is aligned via a closed-form regularized least-squares objective. Crucially, alignment is performed sequentially in topological order — each step uses activations from the partially-merged model, correcting for accumulated upstream errors. After merging in the ambient $d \times d$ space, each operator is factorized back via SVD at a chosen rank $r$, allowing rank expansion ($r > d_k$) as a capacity-control mechanism that retains multi-task singular directions without any training.

SLOA achieves state-of-the-art training-free merging on CLIP/ViT vision benchmarks (8, 14, and 20 tasks) and RoBERTa NLU benchmarks (7 GLUE tasks), outperforming RegMean++ by +6.6 pp on the hardest 20-task vision setting.

Motivation: Why Isolated Merging Fails

Non-invariance of isolated merging

Relative functional error along linear interpolations between equivalent factorizations. Error grows rapidly — equality holds only at $c=1$.

Hidden ambiguity. Any invertible $\mathbf{R}$ gives a different factorization with the same function:

$$(\mathbf{W}_Q\mathbf{R})(\mathbf{W}_K\mathbf{R}^{-\top})^\top = \mathbf{W}_Q\mathbf{W}_K^\top$$

Independent averaging fails. Averaging two such parameterizations yields

$$\bar{\mathbf{M}}_\mathrm{QK} = \tfrac{1}{4}\mathbf{W}_Q(\mathbf{I}+\mathbf{R})(\mathbf{I}+\mathbf{R}^{-1})\mathbf{W}_K^\top \;\neq\; \mathbf{W}_Q\mathbf{W}_K^\top$$

In the scalar case $\mathbf{R}=c\mathbf{I}$, the error grows without bound as $c\to\infty$. Two task models that are functionally identical at the QK level can produce arbitrarily different merged attention scores under isolated weight averaging.

Method

Unified Alignment Objective

For each local operator $\bm{\theta}$, SLOA solves:

$$\bm{\theta}_\mathrm{merge}= \arg\min_{\bm{\theta}}\; \underbrace{\tfrac{1}{T}\!\sum_{t,i}\! \bigl\|\mathcal{A}(\widehat{\mathbf{Z}}_{t,i};\bm{\theta})- \mathcal{A}(\mathbf{Z}_{t,i};\bm{\theta}_t)\bigr\|_F^2}_{\text{activation matching}} \,+\, \underbrace{\tfrac{1}{T}\!\sum_t\!\|\bm{\theta}-\bm{\theta}_t\|^2_{\widehat{\chi}_t}}_{\text{data-aware geometry}}$$

$\widehat{\mathbf{Z}}$ are activations from the partially-merged model; the geometry term uses a Mahalanobis distance that penalizes deviation along directions the calibration activations actually use. For fixed activations the objective is quadratic in $\bm{\theta}$ — closed-form regularized least squares per component.

Local Operators & Metrics

Component Local operator $\bm{\theta}$ Behavior map $\mathcal{A}(\mathbf{Z};\bm{\theta})$ Metric
LayerNorm $(\bm{\gamma},\bm{\beta})$ $\nu(\mathbf{X})\,\mathrm{diag}(\bm{\gamma})+\mathbf{1}\bm{\beta}^\top$ $\mathrm{Cov}(\nu(\chi))$
Linear map $\mathbf{W}$ $\mathbf{X}\mathbf{W}$ $\bm{\Lambda}$
QK component $\tfrac{1}{\sqrt{d_k}}\mathbf{W}_Q\mathbf{W}_K^\top$ $\mathbf{X}\bm{\theta}\mathbf{X}^\top$ $\bm{\Lambda}\otimes\bm{\Lambda}$
VO component $\mathbf{W}_V\mathbf{W}_O$ $\mathbf{A}\mathbf{X}\bm{\theta}$ $\bm{\Lambda}$

$\nu(\mathbf{X})$ = LayerNorm-normalized input; $\mathbf{A}$ = attention matrix; $\bm{\Lambda}=\alpha\,\mathrm{diag}(\mathrm{Cov}(\mathbf{X}))+\epsilon\mathbf{I}$.

SLOA in 4 Steps

1. Collect activations $\{\mathbf{Z}_{t,i}\}$ for each operator from a small calibration set (256 samples per task by default).
2. Order operators topologically: LayerNorm $\to$ QK $\to$ VO $\to$ MLP $W_1$ $\to$ MLP $W_2$.
3. For each operator: forward through the partially-merged model to get $\widehat{\mathbf{Z}}_{t,i}$, then solve closed-form regularized LS.
4. Factorize QK/VO operators via SVD at chosen rank $r_j$ and write back transformer-compatible weights. Rank $r_j > d_k$ retains additional task-specific singular directions.

No gradients, no labels — one deterministic pass through the operator graph.

Results

Vision: ViT-B/32 (CLIP fine-tunes)

Average classification accuracy (%) on 8, 14, and 20 vision tasks. Training-free methods use 256 calibration samples per task.

MethodSamples 8 tasks14 tasks20 tasks
Fine-tuned (oracle)—90.3989.2689.78
Data-Free Weight Averaging—66.3265.3661.10
Task Arithmetic—67.5552.8036.30
TIES-Merging—71.9067.5755.56
TSV-M—83.0778.7973.22
Iso-C—82.5176.5166.54
Iso-CTS—82.7178.3270.22
Training-Free Fisher Merging25670.3566.5062.04
RegMean25682.5077.7572.14
RegMean++25683.9779.5274.06
SLOA ($r{=}64$)25684.8981.9080.08
SLOA ($r{=}128$)25685.5282.6381.02
SLOA ($r{=}256$)256 85.6282.7981.28

SLOA is the best training-free method on all three benchmarks; the gap to baselines widens with more tasks. Rank expansion $r{=}64\to128$ consistently helps. +6.6 pp over RegMean++ on the hardest 20-task setting.

Radar: SLOA vs baselines

Task-wise comparison: SLOA vs. data-free and training-free baselines on ViT-B/32 (20 tasks) and RoBERTa (7 GLUE tasks). SLOA dominates across tasks in both modalities.

Vision: ViT-B/16 (CLIP fine-tunes)

Average classification accuracy (%) on 8, 14, and 20 vision tasks. Training-free methods use 256 calibration samples per task.

MethodSamples 8 tasks14 tasks20 tasks
Fine-tuned (oracle)—92.3491.3191.61
Data-Free Weight Averaging—72.3369.7564.82
Task Arithmetic—77.1460.7939.89
TIES-Merging—77.6071.5160.64
TSV-M—87.1082.1378.31
Iso-C—87.8681.1173.71
Iso-CTS—88.3283.8877.66
Training-Free Fisher Merging25675.5972.0167.01
RegMean25686.2581.5575.50
RegMean++25687.3282.3676.67
SLOA ($r{=}64$)25687.1884.1982.14
SLOA ($r{=}128$)25688.0185.1683.45
SLOA ($r{=}256$)256 88.0985.3083.62

SLOA leads on 14 and 20 tasks; on 8 tasks it is within 0.2 pp of Iso-CTS. The margin over the best baseline grows with task count (+6.95 pp over RegMean++ at 20 tasks).

Vision: ViT-L/14 (CLIP fine-tunes)

Average classification accuracy (%) on 8, 14, and 20 vision tasks. Training-free methods use 256 calibration samples per task.

MethodSamples 8 tasks14 tasks20 tasks
Fine-tuned (oracle)—94.3493.3993.54
Data-Free Weight Averaging—79.8777.5371.15
Task Arithmetic—80.4663.1235.99
TIES-Merging—83.8377.7962.96
TSV-M—90.5788.2183.19
Iso-C—92.1589.3282.80
Iso-CTS—92.7290.0186.12
Training-Free Fisher Merging25682.4175.2669.85
RegMean25690.4587.8281.13
RegMean++25690.9487.8082.26
SLOA ($r{=}64$)25689.3787.5585.67
SLOA ($r{=}128$)25689.9488.2586.38
SLOA ($r{=}256$)256 90.0988.3786.62

On the largest backbone SLOA wins clearly at 20 tasks (+4.36 pp over RegMean++); Iso-CTS edges it on 8/14 tasks, but SLOA is the only method robust as the task count scales up.

NLU: RoBERTa on 7 GLUE Tasks

Per-task metric (Matthews/F1/accuracy) and average on seven GLUE tasks. Training-free methods use 256 calibration samples per task.

MethodSamples CoLASST-2MRPCQQP MNLIQNLIRTEAvg
Individual— 0.6020.9400.8920.914 0.8720.9270.7910.848
Data-Free Weight Averaging— 0.1810.8190.7790.796 0.4380.7110.6170.620
Task Arithmetic— 0.2330.8660.7870.840 0.6370.7300.6100.672
TIES-Merging— 0.2500.8350.7870.852 0.6070.7580.4220.644
TSV-M— 0.3320.8720.8380.860 0.5900.7270.4660.669
Iso-C— 0.3410.8760.8110.799 0.6690.7210.7290.707
Iso-CTS— 0.2940.8200.7070.670 0.4590.6360.6860.610
Training-Free Fisher Merging256 0.1130.8720.7420.781 0.4230.5310.5150.568
RegMean256 0.2800.8930.8070.792 0.7250.8340.6160.707
RegMean++256 0.2450.9140.7700.813 0.7600.8510.6910.721
SLOA256 0.4930.9060.8760.795 0.7210.8560.6810.761
SLOA ($r{=}128$)256 0.5080.9170.8790.814 0.7620.8810.7140.782

bold = best in column; underline = second best. SLOA ($r{=}128$) wins 6/7 tasks and the average.

LLM: LLaMA-2-7B LoRA Experts

Merging three LLaMA-2-7B LoRA experts (math, code, instruction-following) into one model. RoPE blocks a position-independent $\mathbf{W}_Q\mathbf{W}_K^{\top}$, so SLOA is applied to the VO circuit only ($r_{\mathrm{VO}}$ = factorization rank); training-free methods use 256 calibration samples per task.

Method GSM8KHumanEvalIFEvalAvg.
LLaMA-2-7B (pretrained)4.320.0020.898.40
Experts IF-Tulu12.591.2266.1726.66
Code (Magic-Coder)9.7032.9326.9923.21
Math (Meta-Math)65.500.0015.5327.01
Data-Free Weight Averaging40.1118.2940.6733.02
Task Arithmetic58.9815.8542.7039.18
TIES-Merging48.1420.7334.5734.48
TSV-M58.0018.2942.1439.48
Iso-C38.8921.3457.6739.30
Iso-CTS36.0921.3455.4537.63
Training-Free RegMean56.0329.2748.2444.51
RegMean++55.9529.2754.1646.46
SLOA-VO ($r_{\mathrm{VO}}{=}128$)58.6128.0557.8648.17
SLOA-VO ($r_{\mathrm{VO}}{=}256$)59.5127.4458.9648.64
SLOA-VO ($r_{\mathrm{VO}}{=}384$) 59.5127.4460.6349.19

SLOA-VO is the best training-free method, beating RegMean++ by +2.7 avg. pts and every data-free baseline — operator-level merging extends beyond attention-only heads to a 7B decoder-only LLM and to LoRA-adapted experts. Rank expansion ($r_{\mathrm{VO}}{=}128{\to}384$) keeps improving IFEval/avg., mirroring the vision rank-expansion trend.

Ablations

Component ablation

Component ablation on 20-task ViT-B/32. Sequential alignment alone is already strong; adding QK+VO alignment and LayerNorm correction both help.

Rank ablation

Decoupling QK and OV ranks on 20 tasks. QK and OV must be merged jointly — isolating one hurts performance.

Calibration and rank scaling

Performance vs. calibration samples and rank on 20-task ViT-B/32. Rank expansion consistently improves performance; more calibration data helps too.

Alpha sensitivity

Sensitivity to regularization strength $\alpha$ on 8 tasks. Optimal $\alpha$ decreases as more calibration data is available.

NLP calibration scaling

NLP calibration scaling (RoBERTa, $r{=}64$, $\alpha{=}0.25$, 3-seed avg). Performance increases monotonically — 1024 samples (avg 0.799) surpasses $r{=}256$ at 256 samples (0.797).

Key takeaways from ablations:

  • Sequential alignment alone is already strong.
  • QK and OV must be merged jointly — isolating one hurts.
  • Rank expansion consistently improves performance.
  • Optimal $\alpha$ decreases as more calibration data is available.
  • NLP performance increases monotonically with more calibration data — no saturation observed up to 1024 samples.

BibTeX

Citation coming soon.