Multi-Head Attention (MHA). Sum of $H$ heads:
MLP block. Two linear projections separated by a nonlinearity $\sigma(\cdot)$:
Relative functional error along linear interpolations between equivalent factorizations, grows rapidly with the scalar $c$.
SLOA aligns the effective local operators (QK, VO, MLP, LayerNorm), merges them sequentially under partially-merged activations, then factorizes each operator back into transformer-compatible weights.
Unified alignment objective
Local operators & data-aware metrics
| Component | Local operator $\bm{\theta}$ | Behavior map $\mathcal{A}(\mathbf{Z};\bm{\theta})$ | Metric |
|---|---|---|---|
| LayerNorm | $(\bm{\gamma},\bm{\beta})$ | $\nu(\mathbf{X})\,\mathrm{diag}(\bm{\gamma})+\mathbf{1}\bm{\beta}^{\!\top}$ | $\mathrm{Cov}(\nu(\chi))$ |
| Linear map | $\mathbf{W}$ | $\mathbf{X}\mathbf{W}$ | $\bm{\Lambda}$ |
| QK component | $\dfrac{1}{\sqrt{d_k}}\mathbf{W}_Q\mathbf{W}_K^{\!\top}$ | $\mathbf{X}\bm{\theta}\mathbf{X}^{\!\top}$ | $\bm{\Lambda}\otimes\bm{\Lambda}$ |
| VO component | $\mathbf{W}_V\mathbf{W}_O$ | $\mathbf{A}\mathbf{X}\bm{\theta}$ | $\bm{\Lambda}$ |
$\nu(\mathbf{X})$ is the LayerNorm-normalized input; $\mathbf{A}$ is the attention matrix; $\bm{\Lambda}=\alpha\,\mathrm{diag}\bigl(\mathrm{Cov}(\mathbf{X})\bigr)+\epsilon\mathbf{I}$.
After merging the QK / VO operator in $\mathbb{R}^{d\times d}$, SVD the operator $\bar{\bm{\theta}}=\mathbf{U}\bm{\Sigma}\mathbf{V}^{\!\top}$ and factorize at chosen rank $r_j$.
QK circuit (with denominator $c_j=\sqrt{r_j}$):
VO circuit:
No gradients, no labels — one deterministic pass through the operator graph.
Table 1. Average classification accuracy (%) of the merged model under three increasingly difficult settings: 8, 14, and 20 task-specific ViT-B/32 models. Higher is better; training-free methods use 256 calibration examples per task. For SLOA, $r$ is the retained QK/VO factorization rank ($r{=}64$ matches the original head dimension).
| Method | Samples | Avg. Accuracy (%) $\uparrow$ | |||
|---|---|---|---|---|---|
| 8 tasks | 14 tasks | 20 tasks | |||
| Fine-tuned (oracle) | – | 90.38 | 89.25 | 89.77 | |
| Data-Free | Weight Averaging | – | 66.30 | 65.36 | 61.11 |
| Task Arithmetic | – | 67.55 | 52.78 | 60.99 | |
| TIES-Merging | – | 71.88 | 67.57 | 55.57 | |
| TSV-M | – | 83.08 | 78.80 | 73.22 | |
| Iso-C | – | 80.40 | 78.07 | 70.34 | |
| Iso-CTS | – | 82.70 | 78.79 | 73.54 | |
| Training-Free | Fisher Merging | 256 | 69.56 | 67.92 | 62.70 |
| RegMean | 256 | 82.69 | 77.78 | 72.32 | |
| RegMean++ | 256 | 84.51 | 79.91 | 74.54 | |
| SLOA ($r{=}64$) | 256 | 84.90 | 81.78 | 80.17 | |
| SLOA ($r{=}128$) | 256 | 85.47 | 82.54 | 81.14 | |
Table 2. Per-task and average Metric of the merged RoBERTa model on seven GLUE tasks (CoLA, SST-2, MRPC, QQP, MNLI, QNLI, RTE). Higher is better; training-free methods use 256 calibration examples per task. SLOA preserves the original head rank; SLOA ($r{=}128$) expands the QK/VO factorization rank.
| Method | Samples | Metric $\uparrow$ | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| CoLA | SST-2 | MRPC | QQP | MNLI | QNLI | RTE | Avg | |||
| Individual | – | 0.6018 | 0.9404 | 0.8922 | 0.9141 | 0.8720 | 0.9271 | 0.7906 | 0.8483 | |
| Data-Free | Weight Averaging | – | 0.1808 | 0.8188 | 0.7794 | 0.7960 | 0.4383 | 0.7106 | 0.6173 | 0.6202 |
| Task Arithmetic | – | 0.2330 | 0.8658 | 0.7868 | 0.8395 | 0.6371 | 0.7304 | 0.6101 | 0.6718 | |
| TIES-Merging | – | 0.2499 | 0.8349 | 0.7868 | 0.8515 | 0.6072 | 0.7580 | 0.4224 | 0.6444 | |
| TSV-M | – | 0.3324 | 0.8716 | 0.8382 | 0.8598 | 0.5897 | 0.7274 | 0.4657 | 0.6693 | |
| Iso-C | – | 0.3406 | 0.8761 | 0.8108 | 0.7992 | 0.6686 | 0.7210 | 0.7292 | 0.7065 | |
| Iso-CTS | – | 0.2935 | 0.8200 | 0.7070 | 0.6702 | 0.4589 | 0.6363 | 0.6859 | 0.6102 | |
| Training-Free | Fisher Merging | 256 | 0.1126 | 0.8719 | 0.7419 | 0.7808 | 0.4227 | 0.5311 | 0.5150 | 0.5680 |
| RegMean | 256 | 0.2797 | 0.8933 | 0.8069 | 0.7915 | 0.7251 | 0.8336 | 0.6161 | 0.7066 | |
| RegMean++ | 256 | 0.2450 | 0.9136 | 0.7696 | 0.8128 | 0.7603 | 0.8514 | 0.6907 | 0.7205 | |
| SLOA | 256 | 0.4932 | 0.9063 | 0.8764 | 0.7948 | 0.7212 | 0.8555 | 0.6811 | 0.7612 | |
| SLOA ($r{=}128$) | 256 | 0.5080 | 0.9170 | 0.8790 | 0.8144 | 0.7623 | 0.8808 | 0.7136 | 0.7822 | |
bold = best in column; underline = second best. SLOA ($r{=}128$) wins 6/7 tasks and the average.
Component ablation (20 tasks)
Decoupling QK and OV ranks
Scaling with calibration & rank (vision)
Sensitivity to regularization $\alpha$
From left to right: ViT-B/32 on 20 vision tasks vs. data-free / vs. training-free baselines; RoBERTa on 7 GLUE tasks vs. data-free / vs. training-free baselines. SLOA dominates across task in both modalities.
Take-aways
Limitations
Future work