Sequential Local Operator Alignment
for Training-Free Model Merging

Akansh Maurya1   Ya-Wei Eileen Lin2,3   Stefanie Jegelka2,3,4   Sebastian U. Stich1   Rotem Mulayoff1
1CISPA Helmholtz Center for Information Security   2Technical University of Munich   3Munich Center for Machine Learning   4Massachusetts Institute of Technology
Model Merging
Training-Free
Transformer Circuits

TL;DR

  • Existing training-free mergers combine $\mathbf{W}_Q,\mathbf{W}_K,\mathbf{W}_V,\mathbf{W}_O$ independently — ignoring how attention operation actually operattes on inputs.
  • SLOA merges the functional operators $\mathbf{W}_Q\mathbf{W}_K^{\!\top}$ and $\mathbf{W}_V\mathbf{W}_O$ directly.
  • Same operator-alignment principle also handles LayerNorm and the MLP sub-blocks ($\mathbf{W}_1,\mathbf{W}_2$) — SLOA covers the whole transformer, not just attention.
  • Alignment is performed sequentially on activations of the partially-merged model.
  • Re-factorization lets rank act as capacity control.
  • State-of-the-art training-free merging on CLIP/ViT vision and RoBERTa NLU benchmarks.

Model Merging

  • $T$ task-specific models $\{\bm{\Phi}_1,\ldots,\bm{\Phi}_T\}$, each fine-tuned from the same pretrained $\bm{\Phi}_{\text{pre}}$ on its dataset $D_t$.
  • Goal: aggregate them into a single multi-task model without joint retraining on $D_1\cup\cdots\cup D_T$:
$$\begin{gathered} \bm{\Phi}_{\text{merge}} \;=\; g(\bm{\Phi}_1,\ldots,\bm{\Phi}_T) \\[3pt] \text{s.t.}\quad \mathcal{M}_t(\bm{\Phi}_{\text{merge}})\;\approx\;\mathcal{M}_t(\bm{\Phi}_t)\quad\forall\, t. \end{gathered}$$
  • Training-free: a small unlabeled calibration set is allowed; no gradient updates.

Transformer Architecture

Each transformer layer has two sub-modules with residual connections and LayerNorm.

Multi-Head Attention (MHA). Sum of $H$ heads:

$$\mathbf{Z}_{\text{MHA}}=\sum_{j=1}^{H}\mathrm{softmax}\!\Bigl(\tfrac{(\mathbf{X}\mathbf{W}_Q^{(j)})(\mathbf{X}\mathbf{W}_K^{(j)})^{\!\top}}{\sqrt{d_k}}\Bigr)\, \mathbf{X}\,\mathbf{W}_V^{(j)}\mathbf{W}_O^{(j)}$$
  • QK circuit $\mathbf{M}_{\text{QK}}^{(j)}=\mathbf{W}_Q^{(j)}\mathbf{W}_K^{(j)\top}$ — where attention reads.
  • VO circuit $\mathbf{M}_{\text{VO}}^{(j)}=\mathbf{W}_V^{(j)}\mathbf{W}_O^{(j)}$ — what is written.
  • Each circuit is a $d{\times}d$ operator with a hidden rank-$d_k$ bottleneck.

MLP block. Two linear projections separated by a nonlinearity $\sigma(\cdot)$:

$$\mathbf{Z}_{\text{MLP}}=\sigma\!\bigl(\mathbf{Z}\mathbf{W}_1\bigr)\mathbf{W}_2,\qquad \mathbf{W}_1\in\mathbb{R}^{d\times d_{\text{out}}},\;\mathbf{W}_2\in\mathbb{R}^{d_{\text{out}}\times d}.$$
  • SLOA aligns $\mathbf{W}_1$ on the attention output, then $\mathbf{W}_2$ on the merged post-activation $\sigma(\widehat{\mathbf{Z}}\mathbf{W}_1)$ — sequential, closed-form, same template as attention.

Motivation: Isolated merging breaks attention

Non-invariance plot

Relative functional error along linear interpolations between equivalent factorizations, grows rapidly with the scalar $c$.

  • Functional unit. Attention scores depend only on the product $\mathbf{M}_{\mathrm{QK}}=\mathbf{W}_Q\mathbf{W}_K^{\!\top}$ — not on $\mathbf{W}_Q,\mathbf{W}_K$ separately.
  • Hidden ambiguity. Any invertible $\mathbf{R}$ gives a different factorization with the same function:
    $$\bigl(\mathbf{W}_Q\mathbf{R}\bigr)\bigl(\mathbf{W}_K\mathbf{R}^{-\top}\bigr)^{\!\top} \;=\;\mathbf{W}_Q\mathbf{W}_K^{\!\top}.$$
  • Independent averaging fails. Averaging the two parameterizations element-wise gives $\bar{\mathbf{W}}_Q=\tfrac{1}{2}\mathbf{W}_Q(\mathbf{I}+\mathbf{R})$, $\bar{\mathbf{W}}_K=\tfrac{1}{2}\mathbf{W}_K(\mathbf{I}+\mathbf{R}^{-\top})$, so
    $$\bar{\mathbf{M}}_{\mathrm{QK}}=\bar{\mathbf{W}}_Q\bar{\mathbf{W}}_K^{\!\top} =\tfrac{1}{4}\mathbf{W}_Q(\mathbf{I}+\mathbf{R})(\mathbf{I}+\mathbf{R}^{-1})\mathbf{W}_K^{\!\top} \;\neq\;\mathbf{W}_Q\mathbf{W}_K^{\!\top}.$$
  • Unbounded blow-up. Scalar case $\mathbf{R}=c\mathbf{I}$:
    $$\bar{\mathbf{M}}_{\mathrm{QK}}=\dfrac{(1+c)(1+c^{-1})}{4}\,\mathbf{M}_{\mathrm{QK}} \;\;\xrightarrow[c\to\infty]{}\;\;\infty.$$
    Equality holds only at $c=1$; error grows without bound otherwise.
  • Consequence. Two task models that are functionally identical at the QK level can produce arbitrarily different merged attention scores under isolated weight averaging.

Key idea

  • Merge at the level of local functional operators, not raw weight tensors.
  • Each component's behavior is linear in its operator $\bm{\theta}$ — closed-form least squares.
  • Apply alignment sequentially: each merge corrects accumulated upstream error.
  • Factorize the merged operator back; the chosen rank controls multi-task capacity.

Method: Sequential Local Operator Alignment

SLOA illustration

SLOA aligns the effective local operators (QK, VO, MLP, LayerNorm), merges them sequentially under partially-merged activations, then factorizes each operator back into transformer-compatible weights.

Unified alignment objective

$$\bm{\theta}_{\mathrm{merge}}= \arg\min_{\bm{\theta}}\; \underbrace{\tfrac{1}{T}\!\sum_{t,i}\! \bigl\|\mathcal{A}(\widehat{\mathbf{Z}}_{t,i};\bm{\theta})- \mathcal{A}(\mathbf{Z}_{t,i};\bm{\theta}_t)\bigr\|_F^2}_{\text{activation matching}} \,+\, \underbrace{\tfrac{1}{T}\!\sum_t\!\|\bm{\theta}-\bm{\theta}_t\|^2_{\widehat{\chi}_t}}_{\text{data-aware geometry}}$$

Local operators & data-aware metrics

Component Local operator $\bm{\theta}$ Behavior map $\mathcal{A}(\mathbf{Z};\bm{\theta})$ Metric
LayerNorm $(\bm{\gamma},\bm{\beta})$ $\nu(\mathbf{X})\,\mathrm{diag}(\bm{\gamma})+\mathbf{1}\bm{\beta}^{\!\top}$ $\mathrm{Cov}(\nu(\chi))$
Linear map $\mathbf{W}$ $\mathbf{X}\mathbf{W}$ $\bm{\Lambda}$
QK component $\dfrac{1}{\sqrt{d_k}}\mathbf{W}_Q\mathbf{W}_K^{\!\top}$ $\mathbf{X}\bm{\theta}\mathbf{X}^{\!\top}$ $\bm{\Lambda}\otimes\bm{\Lambda}$
VO component $\mathbf{W}_V\mathbf{W}_O$ $\mathbf{A}\mathbf{X}\bm{\theta}$ $\bm{\Lambda}$

$\nu(\mathbf{X})$ is the LayerNorm-normalized input; $\mathbf{A}$ is the attention matrix; $\bm{\Lambda}=\alpha\,\mathrm{diag}\bigl(\mathrm{Cov}(\mathbf{X})\bigr)+\epsilon\mathbf{I}$.

  • Activation term matches local input–output behavior on calibration data.
  • Geometry term is a Mahalanobis distance: penalizes deviation along directions the calibration activations actually use.
  • For fixed activations the objective is quadratic in $\bm{\theta}$ — closed-form regularized least squares per component.
  • Sequential: downstream operators are aligned on activations from the partially merged model $\widehat{\mathbf{Z}}$.

Rank as capacity control

After merging the QK / VO operator in $\mathbb{R}^{d\times d}$, SVD the operator $\bar{\bm{\theta}}=\mathbf{U}\bm{\Sigma}\mathbf{V}^{\!\top}$ and factorize at chosen rank $r_j$.

QK circuit (with denominator $c_j=\sqrt{r_j}$):

$$\mathbf{W}_Q^{(j)}=\mathbf{U}_{r_j}(c_j\bm{\Sigma}_{r_j})^{1/2},\quad \mathbf{W}_K^{(j)}=\mathbf{V}_{r_j}(c_j\bm{\Sigma}_{r_j})^{1/2}$$

VO circuit:

$$\mathbf{W}_V^{(j)}=\mathbf{U}_{r_j}\bm{\Sigma}_{r_j}^{1/2},\qquad \mathbf{W}_O^{(j)}=\bm{\Sigma}_{r_j}^{1/2}\mathbf{V}_{r_j}^{\!\top}$$
  • $r_j = d_k$ preserves the original head width.
  • $r_j > d_k$ retains additional task-specific singular directions.
  • The sum of task heads can span a larger subspace than any single rank-$d_k$ head — expansion avoids unnecessary truncation.

SLOA in 4 steps

  1. Collect activations $\{\mathbf{Z}_{t,i}\}$ for each operator from a small calibration set.
  2. Order operators topologically $o_1\!\prec\!\cdots\!\prec\!o_M$ (LayerNorm $\to$ QK $\to$ VO $\to$ MLP).
  3. For each $o_m$: forward through the partially-merged model to get $\widehat{\mathbf{Z}}_{t,i}$, then solve closed-form regularized LS.
  4. Factorize QK/VO via SVD at chosen rank $r_j$ and write back transformer-compatible weights.

No gradients, no labels — one deterministic pass through the operator graph.

Vision: Merging ViT-B/32 (CLIP fine-tunes)

Table 1. Average classification accuracy (%) of the merged model under three increasingly difficult settings: 8, 14, and 20 task-specific ViT-B/32 models. Higher is better; training-free methods use 256 calibration examples per task. For SLOA, $r$ is the retained QK/VO factorization rank ($r{=}64$ matches the original head dimension).

MethodSamples Avg. Accuracy (%) $\uparrow$
8 tasks14 tasks20 tasks
Fine-tuned (oracle)–90.3889.2589.77
Data-FreeWeight Averaging–66.3065.3661.11
Task Arithmetic–67.5552.7860.99
TIES-Merging–71.8867.5755.57
TSV-M–83.0878.8073.22
Iso-C–80.4078.0770.34
Iso-CTS–82.7078.7973.54
Training-FreeFisher Merging25669.5667.9262.70
RegMean25682.6977.7872.32
RegMean++25684.5179.9174.54
SLOA ($r{=}64$)25684.9081.7880.17
SLOA ($r{=}128$)25685.4782.5481.14
  • SLOA is the best training-free method on all three benchmarks; the gap to baselines widens with more tasks.
  • +6.6 pts over RegMean++ on the hardest 20-task setting.
  • Rank expansion $r{=}64{\to}128$ consistently helps — the merged operator's spectrum spans a larger subspace than any single head.

NLU: Merging RoBERTa on 7 GLUE tasks

Table 2. Per-task and average Metric of the merged RoBERTa model on seven GLUE tasks (CoLA, SST-2, MRPC, QQP, MNLI, QNLI, RTE). Higher is better; training-free methods use 256 calibration examples per task. SLOA preserves the original head rank; SLOA ($r{=}128$) expands the QK/VO factorization rank.

MethodSamplesMetric $\uparrow$
CoLASST-2MRPCQQPMNLIQNLIRTEAvg
Individual–0.60180.94040.89220.91410.87200.92710.79060.8483
Data-FreeWeight Averaging–0.18080.81880.77940.79600.43830.71060.61730.6202
Task Arithmetic–0.23300.86580.78680.83950.63710.73040.61010.6718
TIES-Merging–0.24990.83490.78680.85150.60720.75800.42240.6444
TSV-M–0.33240.87160.83820.85980.58970.72740.46570.6693
Iso-C–0.34060.87610.81080.79920.66860.72100.72920.7065
Iso-CTS–0.29350.82000.70700.67020.45890.63630.68590.6102
Training-FreeFisher Merging2560.11260.87190.74190.78080.42270.53110.51500.5680
RegMean2560.27970.89330.80690.79150.72510.83360.61610.7066
RegMean++2560.24500.91360.76960.81280.76030.85140.69070.7205
SLOA2560.49320.90630.87640.79480.72120.85550.68110.7612
SLOA ($r{=}128$)2560.50800.91700.87900.81440.76230.88080.71360.7822

bold = best in column; underline = second best. SLOA ($r{=}128$) wins 6/7 tasks and the average.

Ablations

Component ablation (20 tasks)

Decoupling QK and OV ranks

Scaling with calibration & rank (vision)

Sensitivity to regularization $\alpha$

  • Sequential alignment alone is already strong.
  • QK and OV must be merged jointly — isolating one hurts.
  • Rank expansion consistently improves performance.
  • Optimal $\alpha$ decreases as more calibration data is available.

Task-wise comparison: SLOA vs. data-free and training-free baselines

Radar plots: SLOA vs baselines

From left to right: ViT-B/32 on 20 vision tasks vs. data-free / vs. training-free baselines; RoBERTa on 7 GLUE tasks vs. data-free / vs. training-free baselines. SLOA dominates across task in both modalities.

Conclusions & Limitations

Take-aways

  • Treat $\mathbf{W}_Q\mathbf{W}_K^{\!\top}$ and $\mathbf{W}_V\mathbf{W}_O$ as the units to merge — not the raw matrices.
  • Sequential, data-aware alignment removes parameterization ambiguity and curbs cross-layer error accumulation.
  • Rank expansion grows multi-task capacity without training.
  • New SOTA training-free merging on CLIP/ViT vision and RoBERTa NLU.

Limitations

  • Needs calibration data — more than purely data-free baselines require.
  • Higher offline merging cost: $\sim$14 min vs. $\sim$2 min for Fisher/RegMean on 20-task ViT-B/32 — strictly one-time, no backprop.
  • Rank expansion increases inference cost: per-head $O(L^2 r_j)$ when $r_j>d_k$; trade-off between capacity and latency.

Future work

  • Reduce calibration / merging cost and automatically select per-head ranks under deployment budgets.
  • Extend operator-level merging to decoder-only and encoder-decoder transformers.
  • Adapt to parameter-efficient fine-tunes (LoRA, adapters) where weights live in low-rank deltas.