Paper

How to Hand Off KV Caches As-Is Between Different LLMs: HeteroFold

TL;DR: HeteroFold, a method that reuses the KV cache between models with different tokenizers and architectures without the receiver re-prefilling, beats existing cache transfer in all six directions and is 10.74× faster than native prefill at 32K context (source: §4.1, Tab.1, Tab.4).

Core Idea

In heterogeneous multi-agent settings, pass the KV cache directly instead of text. Freeze the models, and fold only the token, layer, and distribution mismatches into external affine maps (source: §3, Fig.3).

HeteroFold’s claim is simple. If the KV already computed by the sender can be transformed into the receiver’s space, the receiver does not need to re-prefill the long shared context. The point is not to make the values similar, but to correct them so that the receiver’s attention and output match the native ones (source: §1, Fig.2).

Background: The Problem They Solved

The Research Gap

Recent multi-agent LLMs use different model families for different roles. For example, X-MAS assigns different families to the solver, evaluator, and aggregator (source: §1).

The problem is communication. Existing text-based communication (TextMas) has the receiver re-prefill the shared text and rebuild the KV. As the context grows longer, this redundant computation becomes the main cause of latency (source: §1).

The state of the art at the time was split as follows (source: §2):

  • Cross-family communication exists, but it requires prefill: TextMas is lossless, but receiver prefill is mandatory. C2C requires a receiver-built prompt cache, so it does not eliminate prefill either. RECURSIVEMAS passes hidden representations, but they are not decodable KV and must pass through the receiver’s layers again.
  • Prefill-free exists, but only within the same family: LATENTMAS assumes cache compatibility. Dense Latent and KV Ridge map the sender cache to the receiver cache and avoid prefill, but they were evaluated only within the same tokenizer and similar architectures.

In other words, no cross-family method simultaneously handled tokenizer mismatch, depth mismatch, and KV representation mismatch while completely eliminating receiver prefill. The authors present this as the first prefill-free transfer that explicitly includes cross-tokenizer alignment (source: §1-§2).

Central Hypothesis

The authors hypothesize that by aligning tokens to shared character boundaries, matching multi-layer sender features to receiver statistics via recoloring, and folding corrections computed against receiver attention and output into fixed affine maps, one can eliminate receiver prefill for cross-family KV reuse and preserve downstream behavior while keeping both models frozen (source: §3).

Three Original Contributions

  1. A new alignment technique: Token Alignment (TA): A method that maps tokens from different tokenizers to the same character end position. Unmatched receiver positions reuse the sender state at the most recent shared boundary and are not averaged (source: §3.1, Appx.B.1). This constitutes a new preprocessing and matching technique.
  2. A new architecture component: multi-layer, cross-head Recolor mapping: Concatenate three nearby sender layers matched by proportional depth, unfold the heads, then match the receiver mean and covariance via whitening–Procrustes–recoloring. A separate affine map is built per receiver layer for K and V (source: §3.1-§3.2). This constitutes a new mapping module.
  3. A new training technique, purpose-built: output-aware calibration and folding: Output-aware correction for KV quantization is adapted for cross-family transfer with separate K/V objectives, training only rank-16 low-rank corrections $B$, $G$ and then folding them in the form $A_{final}=A(I+BG)$ so that no extra module exists at inference (source: §3.3). This constitutes a new application of an existing methodology and a new training objective.

Strengths from the Authors’ Perspective

The authors’ argument is that “low reconstruction error does not mean the receiver behaves well.” On the Qasper average, KV Ridge+TA has a cache reconstruction error of 0.32, lower than HeteroFold’s 0.43, but it loses badly on attention-weight KL (1.14 vs. 0.51) and attention output error (0.72 vs. 0.40) (source: §1, Fig.2).

They therefore argue that distribution moment matching and receiver query, attention, and output matching are needed, not token-wise $l_2$ reconstruction. Indeed, KV Ridge with only TA added jumps substantially (e.g., Qwen3-4B→Llama-3.1-8B HotpotQA F1 from 0.74 to 34.04), but HeteroFold leads further at 38.48. This is interpreted as the effect of K/V mapping and correction beyond token correspondence (source: §4.1, Tab.1).

The overall HeteroFold flow: token and layer alignment, recoloring, output-aware calibration, and folding

The New Approach: HeteroFold

Let us set notation. $S$ is the frozen sender, $R$ is the frozen receiver, $L_S$, $L_R$ are the layer counts, $C_S(x)$ is the sender cache for context $x$, and $\hat{C}_R(x)$ is the transformed receiver-compatible cache. $c \in \{K,V\}$ is the key/value role, and $d^{KV}$ is the unfolded KV feature dimension (source: §3).

HeteroFold maps K at the point before key normalization and RoPE, and V at the point after projection. The receiver then applies native key normalization and RoPE exactly once (source: §3, Appx.B.1).

There are four steps.

Step 1: Token Alignment (TA)
Tokenize the same raw text with each tokenizer, and record the character end offsets $e^S_i$, $e^R_j$ of non-empty tokens. Pairs that share the same end position are grouped, and if there is no shared point, the sender state at the most recent shared boundary is filled forward. The segment before the first match uses the first sender token (source: §3.1, Appx.B.1).

Step 2: Layer Alignment (LA)
For receiver layer $l$, find the sender center layer by proportional depth and group the surrounding 4-offset neighbors.

$$ \pi(l) = 1 + \text{round}\left((l-1)\frac{L_S-1}{L_R-1}\right),\quad \mathcal{N}(l)=\text{clip}_{[1,L_S]}(\pi(l)-4,\pi(l),\pi(l)+4) $$

For each aligned token pair $(i_n,j_n)$, concatenate the sender’s three layer features: $x^{l,c}_n=[c^{S,t_1}_{i_n} \| c^{S,t_2}_{i_n} \| c^{S,t_3}_{i_n}]$, and the target is $y^{l,c}_n=c^{R,l}_{j_n}$. Because the heads are unfolded, this becomes a cross-head mapping even when the head counts differ (source: §3.1).

Step 3: Recolor
For $X \in \mathbb{R}^{N \times \tilde{d}_S}$, $Y \in \mathbb{R}^{N \times d_R}$ collected from calibration prompts, scale by the scalar RMS $r$ and set $Z=X/r$. Compute the weighted means $\mu_Z$, $\mu_Y$, the covariances $\Sigma_Z$, $\Sigma_Y$, and the cross-covariance $\Sigma_{ZY}$, then find the Procrustes rotation $R^{\star}=UQ^{\top}$ after whitening (source: §3.2).

$$ \Sigma_Z^{-1/2}\Sigma_{ZY}\Sigma_Y^{-1/2}=U\Lambda Q^{\top} $$$$ A=\Sigma_Z^{-1/2}R^{\star}\Sigma_Y^{1/2},\quad b=\mu_Y-\mu_Z A,\quad y^{(0)}=(x/r)A+b $$

Ideally this satisfies $\mu_Z A+b=\mu_Y$, $A^{\top}\Sigma_Z A=\Sigma_Y$. With numerical stabilization, it is an approximate match in practice (source: §3.2).

Step 4: Output-aware Calibration and Folding
For fixed receiver queries, split the difference between the native and transformed attention outputs into key-induced and value-induced changes.

$$ \hat{O}-O^{\star}=\underbrace{[(\hat{P}-P^{\star})V^{\star}]W_O^{\top}}_{\text{key-induced change}}+\underbrace{[\hat{P}(\hat{V}-V^{\star})]W_O^{\top}}_{\text{value-induced change}} $$

K optimizes the attention-pattern KL and the post-output-projection error, while V optimizes the final output norm error under the fixed $\hat{P}$. In the rank $\rho_c=16$ correction $\hat{y}=y^{(0)}+(y^{(0)}B)G$, only $B$ and $G$ are trained for 4 epochs, and $G$ is initialized to zero. Only query positions in the question span are used, and neither answers nor generation trajectories are used (source: §3.3, Appx.B.2).

At inference, it is folded.

$$ A^{final}=A(I+BG),\quad b^{final}=b(I+BG),\quad \Phi_{l,c}(x_j)=(x_j/r)A^{final}+b^{final} $$

Only one fixed affine map remains per layer and role, with no separate correction module (source: §3.3).

Motivation experiment showing that multi-layer input gives higher receiver KV prediction $R^2$ than a single layer, that recoloring matches the distribution, and that correction reduces attention output error

How It Works: A Concrete Example Walkthrough

Take the toy example “The unbelievable result.” for a graduate student. Counting character positions, The ends at 3, unbelievable at 16, result at 23, and . at 24 (source: §3.1).

  • Qwen3 tokenization: The | unbelievable | result | . → end positions 3, 16, 23, 24
  • Ministral-3 tokenization: The | unbel | iev | able | result | . → end positions 3, 9, 12, 16, 23, 24

In the Qwen3→Ministral-3 direction, receiver boundaries 9 and 12 have no shared boundary, so the Qwen state at the most recent shared boundary 3 is reused. In the reverse direction, the Qwen token ending at character 16 uses the Ministral state ending at the same 16 as-is. The point is to match causal prefix endpoints rather than averaging the intermediate unbel, iev states (source: §3.1).

Layer alignment also has a numerical example. If sender $L_S=32$, receiver $L_R=28$, and receiver $l=14$, then $\pi(14)=1+\text{round}(13 \times 31/27)\approx 16$. The neighbors become $(12,16,20)$, and the unfolded K features of the three layers are concatenated into a 3× dimension input $x$. This is matched by Recolor to the K target $y$ at receiver layer 14 (source: §3.1).

Why this design is necessary becomes clear from the failure of same-index pairing. Grouping the same $i$-th tokens points to different character positions because the tokenizers differ. Across 48 HotpotQA documents, the average character position mismatch grows to hundreds of characters as the context lengthens, whereas TA stays at 0.003 to 0.20 characters (source: §5.2, Fig.5).

Token alignment error showing that same-index pairing drifts by hundreds of characters as context lengthens, while shared-boundary-based TA stays below 0.2 characters

The Secret Weapon: What Happens Without TA

Consider the Ministral-3-14B→Llama-3.1-8B direction ablation. All variants use the same calibration data (source: §4.3, Tab.3).

VariantARC-C Accuracy (%)GSM8K Accuracy (%)Qasper F1HotpotQA F1QuALITY Accuracy (%)
HeteroFold full90.3663.4633.7250.2979.00
Same-index pairing69.207.133.190.7224.59
Single-layer mapping82.0046.1722.3437.8771.52
Head-local mapping71.8426.699.7117.3854.79
TA+Ridge87.034.553.276.8247.41
TA+Recolor (no correction)90.1962.7726.5240.7179.34

The largest drop comes from removing TA. HotpotQA F1 evaporates from 50.29 to 0.72, about 49.57%p. The mechanism is clear. If the token correspondence is off, then no matter how good the subsequent linear map is, it learns from different prefixes. The collapse of head-local mapping (Qasper 33.72→9.71) is in the same vein. Because the sender and receiver split heads differently, mapping within a head alone cannot carry cross-head information (source: §4.3, Tab.3).

The Ridge vs. Recolor comparison is also instructive. TA+Ridge collapses to GSM8K 4.55% and Qasper 3.27. Adding output-aware correction recovers it to 55.57% and 21.65%, but it still falls short of TA+Recolor+correction (63.46%, 33.72%). Ridge fixates on token-wise $l_2$, which skews the distribution tails and attention logit scale, whereas Recolor first matches the receiver mean and covariance, thereby preserving the attention input distribution. The fact that key correction yields a larger gain in long context is consistent with attention patterns dominating in long contexts (source: §4.3, Tab.3, Fig.4).

Performance Validation: Main Results

Experimental Setup at a Glance

  • Models: meta-llama/Llama-3.1-8B-Instruct, Qwen/Qwen3-4B, mistralai/Ministral-3-14B-Instruct-2512-BF16, all frozen in BF16, with Qwen in non-thinking mode (source: Appx.A.1).
  • Tokenization: process the same prompt with each tokenizer and handle chat template tokens separately (source: Appx.A.1).
  • Calibration: 1,600 training examples (800 Open-R1 + 800 HotpotQA training split), 400 held out, excluding answers and solutions, non-overlapping with evaluation. KV Ridge is reimplemented with $k=8$, ridge 0.01, and Dense Latent with a 2-layer GELU mapping plus a 2-stage training on 512-token receiver traces (source: §4, Appx.A.3, Appx.B.3).
  • Evaluation: 4 long-context benchmarks (Qasper 200, HotpotQA 200, LoCoMo 1,540, QuALITY 2,086) and 5 short-context benchmarks (ARC-C 1,172, MMLU 14,042, WinoGrande 1,267, HellaSwag 10,042, GSM8K 1,319), plus 65 HIDDENBENCH tasks. Qasper and HotpotQA use temperature 0.6, top-p 0.95, top-k 20; LoCoMo and GSM8K use greedy decoding (source: Appx.A.1-A.2).
  • Transformer configuration: GQA-based KV head sharing, RoPE for positions, RoPE applied after receiver key normalization, RoPE computations in FP32 (source: §3.3, Appx.B.1).

Key Result 1: First Place on All Long-Context Tasks Across 6 Cross-Family QA Directions

HeteroFold achieves the best cache-transfer performance on all 4 long-context benchmarks in all 6 directions. This is unusual given that it has no receiver prefill (source: §4.1, Tab.1).

A few representative numbers follow. Units are F1 or accuracy (%).

  • Qwen3-4B→Llama-3.1-8B: Qasper 29.12, HotpotQA 38.48, LoCoMo 29.66, QuALITY 62.85% — all above the runner-up (KV Ridge+TA 16.06, 34.04, 13.56, 54.60%) (source: Tab.1).
  • Ministral-3-14B→Qwen3-4B: Qasper 42.00, HotpotQA 52.74, LoCoMo 35.59, approaching the text ceiling of 43.07, 56.05, 38.92 (source: Tab.1).
  • Llama-3.1-8B→Ministral-3-14B: Qasper 21.58 vs. 11.41, HotpotQA 39.00 vs. 13.33, GSM8K 72.33% vs. 47.99%, roughly 1.5–2.9× over KV Ridge+TA (source: Tab.1).

Short context is mostly ahead as well. In Ministral-3-14B→Llama-3.1-8B, it leads the baselines with ARC-C 90.36%, GSM8K 63.46%, QuALITY 79.00% (source: Tab.1).

Receiver behavior preservation showing HeteroFold is closest to native on next-token KL, answer agreement, and attention KL relative to native

Key Result 2: On Equal Footing with Text Communication on Multi-Agent HIDDENBENCH

After 15 rounds of communication, average accuracy is 30.6%, majority-vote accuracy 29.2%, and invalid rate 3.2%, matching or slightly ahead of TextMas’s 29.2%, 27.7%, 4.0%. The gap to Dense Latent+TA 18.5% and KV Ridge+TA 21.8% is large, and contrasts with KV Ridge+TA’s invalid rate of 20.9% (source: §4.2, Tab.2).

The significance is that collective decision-making is preserved even when only the KV of newly generated messages is mapped, beyond single-hop QA. It also holds in the setting where Qwen→Qwen stays as text and only the rest is cache-transferred (source: Appx.A.2).

Key Result 3: Latency

This is on 2×H100 80GB NVLink with FlashAttention-2, batch size 2, and 4K/16K/32K QuALITY contexts. Excluding sender prefill, model loading, and offline fitting, the median is the sum of tokenization+TA, inter-GPU payload copy, K/V map + cache construction, and receiver processing up to the first token (source: §5.1, Appx.E.1, Tab.4, Tab.I).

  • Llama-3.1-8B→Ministral-3-14B at 32,768 tokens: HeteroFold 481.3 ms, native prefill 5167.4 ms → 10.74× reduction. It is 1.18× faster than KV Ridge+TA 569.9 ms and 1.47× faster than Dense Latent+TA 705.7 ms.
  • Qwen3-4B→Ministral-3-14B at 32,768 tokens: 483.2 ms vs. 5178.2 ms → 10.72×, ahead of the baselines 587.7 ms and 720.5 ms.
  • At 4K it is 3.47–3.75×, and at 16K it is 7.24–8.11×, with the gain growing as context lengthens (source: Tab.4).

The main source of the gain is the smaller K/V map + cache construction step. HeteroFold uses a single affine stage and three sender layers per receiver layer, while Dense Latent uses a 2-layer MLP and KV Ridge uses 8 layers (source: §5.1, Fig.C).

What the authors emphasize most as evidence of success is first place on all long-context tasks, the 10.7× latency reduction, and the text-parity performance on HIDDENBENCH (source: §4.1-§4.2, §5.1).

Critical Comparison: Where It Wins and Where It Slips

The strongest comparison point is the K/V design after sharing TA. Even when pitted against Dense Latent+TA and KV Ridge+TA, which are given the same TA, HeteroFold is closest to native on next-token KL, answer agreement, and attention KL on QuALITY (source: §5.2, Fig.6).

It also holds first place on all long-context tasks within the same family (4 directions of Qwen3 4B/8B/14B), showing generality beyond tokenizer alignment. The fact that KV Ridge remains competitive on some short-context tasks within the same family is evidence that the baseline reimplementation is on track (source: Appx.C, Tab.G).

Conversely, there are points it loses. In some short-context tasks where the receiver is Ministral or Qwen, it loses to KV Ridge+TA. Examples: Ministral-3-14B→Qwen3-4B WinoGrande 53.83% vs. 54.78%, HellaSwag 41.46% vs. 45.12%. Llama-3.1-8B→Qwen3-4B HellaSwag also loses, 41.74% vs. 44.54% (source: Tab.1). A plausible interpretation is that Ridge’s token-wise fit happens to work well on short multiple-choice tasks, or that HeteroFold’s covariance matching can be excessive for short distributions. Also, GSM8K question-reconstruction ordered overlap is HeteroFold 83.04%, KV Ridge+TA 76.68%, Dense Latent+TA 39.74%, ahead on lexical reconstruction, but exact match is only 16/100, showing that lexical preservation and task accuracy are not one-to-one (source: Appx.D.2, Tab.H).

Our Perspective: Strengths, Limitations, and Why This Research Matters

The strength is the way the problem is decomposed. The responsibilities of each layer are clear: token correspondence (TA), depth aggregation (LA+3 layers), distribution matching (Recolor), behavior matching (output-aware separate K/V objective), and cost elimination (folding). In particular, calibrating keys and values separately in line with the decomposition of the attention equation is a spot where the KV quantization literature is brought in but redesigned for the transfer problem (source: §3.3).

It is also convincing from a systems standpoint. The storage map parameters are 201M–252M per direction, 0.40GB–0.50GB in BF16, about 25% of Dense Latent and about 62% less than KV Ridge. The construction cost is about 2 H100 GPU-hours per direction (source: Appx.B.4, Tab.F). Across all six directions, both latency and extra peak memory are lower than KV Ridge+TA (source: Appx.E.3, Fig.D).

The limitations are clear too. They are not stated by the authors, but they emerge from the analysis.

  • The maps are large and direction-dependent. Six directions require up to about 1.5B parameters’ worth of maps, and each new model pair needs 1,600+400 prompt calibration again (source: Appx.B.3-B.4).
  • The transfer latency comparison assumes the sender has already processed the context. Otherwise one must add sender prefill+capture of 280 ms, 1328 ms, 3381 ms for Llama (4K/16K/32K, 2-context batch) and 235 ms, 1163 ms, 3083 ms for Qwen (source: Appx.E.4).
  • Mixing Open-R1 and HotpotQA is needed for calibration to perform. Using a single corpus degrades ARC-C, HotpotQA, and QuALITY alike. That is, recalibration may be needed for out-of-distribution contexts (source: Appx.B.3, Tab.E).
  • There is sensitivity to rank and layer count. Rank 0 (Recolor only) drops to HotpotQA 30.47, and using only 1 layer collapses to 22.52. Three and five layers are similar, so three was chosen, but the optimal neighborhood may differ when the model pair changes (source: Appx.B.3, Tab.C-D).
  • As a generalization limitation, the evaluation focuses on three models (Llama/Qwen/Ministral) and English QA and HIDDENBENCH. Stability over long decoding trajectories was verified only in GSM8K teacher-forced replay (source: Appx.A.1, Appx.D.3).

Even so, the reason this research matters is that as heterogeneous agents become the norm, “who reads the long context how many times” dominates cost. If text passing repeats prefill for each agent, cache passing shares what has been read once. HeteroFold is the first practical bridge that makes that sharing possible across different vocabularies and depths (source: §6).

What’s Next?: The Road Ahead

In the conclusion, the authors only state that they will move toward efficient cache sharing among heterogeneous agents and a more general cross-family KV interface, without enumerating concrete follow-ups (source: §6).

In light of the limitations, reasonable next steps are as follows.

  • Direction-agnostic, lightweight maps: There is a need to explore compressing the 0.4GB–0.5GB map with a hypernetwork or low-rank factorization, and a general-purpose adapter that works without retraining when the model pair changes (source: Appx.B.4).
  • Online, incremental calibration: In environments where messages are generated, such as HIDDENBENCH, the pre-fitted map and the arriving message distribution can diverge. Online correction that lightly updates $B$, $G$ from receiver output feedback is natural (source: §4.2, Appx.A.2).
  • Optimization including sender cost: In practice the sender prefill can be the bottleneck, so end-to-end latency and memory (including payload copy) across the sender-transfer-receiver pipeline should be optimized together (source: Appx.E.3-E.4).
  • Extension to longer generation, multilingual, and modalities: Since calibration used only prompt query positions, extension to thousands of generated tokens, non-English tokenizers, and VLM hidden states is a validation task (source: §3.3, Appx.D.3).

From an era when different models had to re-read the same text every time to explain it to each other, we can now hand over the traces of having read it. HeteroFold has, in effect, produced both a dictionary and a grammar for translating those traces.

Tables from the paper

Tables converted mechanically from the arXiv e-print LaTeX source. The numbers are the paper’s own and did not pass through a model.

Table 1. Cross-family transfer results across six directions. Dense Latent and KV Ridge use same-index token pairing, while +TA replaces only token correspondence with our shared character-boundary alignment. TextMas uses lossless text communication with native receiver prefill. Higher is better. Bold indicates the best cache-transfer result within each direction and benchmark; TextMas is excluded from emphasis.

Transfer directionMethodLong Context QA QasperLong Context QA HotpotQALong Context QA LoCoMoLong Context QA QuALITYShort Context QA ARC-CShort Context QA MMLUShort Context QA WinoGrandeShort Context QA HellaSwagShort Context QA GSM8K
(F1 $\uparrow$)(F1 $\uparrow$)(F1 $\uparrow$)(Acc. $\uparrow$)(Acc. $\uparrow$)(Acc. $\uparrow$)(Acc. $\uparrow$)(Acc. $\uparrow$)(Acc. $\uparrow$)
Qwen3-4B $\rightarrow$ Llama-3.1-8BTextMas44.6457.5452.3375.0782.7666.7465.8270.7885.67
Dense Latent2.830.600.6523.1524.7424.1150.2848.670.91
Dense Latent + TA3.171.371.6826.4627.6526.5650.7550.174.09
KV Ridge2.780.741.4125.1724.4030.3650.3651.400.53
KV Ridge + TA16.0634.0413.5654.6076.7961.4252.4959.5743.21
HeteroFold (Ours)29.1238.4829.6662.8578.9263.2857.3065.6055.12
Ministral-3-14B $\rightarrow$ Qwen3-4BTextMas43.0756.0538.9270.6187.8067.9555.0941.5490.14
Dense Latent2.560.220.2224.9825.0027.9951.2244.661.82
Dense Latent + TA3.661.320.4325.8926.3731.4649.5744.3415.16
KV Ridge2.771.451.7325.4142.1544.2750.9937.457.05
KV Ridge + TA39.0248.2414.7173.2088.1471.7354.7845.1287.11
HeteroFold (Ours)42.0052.7435.5975.5588.9972.1453.8341.4689.23
Llama-3.1-8B $\rightarrow$ Qwen3-4BTextMas43.0756.0538.9270.6187.8067.9555.0941.5490.14
Dense Latent2.721.400.6725.3124.4927.0050.6742.960.76
Dense Latent + TA4.351.390.5825.5025.0027.2451.6243.536.52
KV Ridge2.880.941.0026.9426.6231.0950.4336.931.59
KV Ridge + TA31.6325.1221.9159.5974.3261.2551.3844.5470.66
HeteroFold (Ours)35.9949.4434.8362.9077.0562.3952.8841.7470.74
Qwen3-4B $\rightarrow$ Ministral-3-14BTextMas45.5964.5652.0183.3292.6675.6361.5668.3394.69
Dense Latent3.191.661.2326.5136.2632.0750.0450.623.87
Dense Latent + TA5.074.263.3738.9742.6641.3150.2050.1428.81
KV Ridge2.170.441.4327.8526.4526.7851.1452.471.74
KV Ridge + TA7.7517.099.3362.9484.1363.4153.1261.7460.20
HeteroFold (Ours)20.5742.6517.7566.3085.4966.2455.8865.2277.26
Ministral-3-14B $\rightarrow$ Llama-3.1-8BTextMas44.6457.5452.3375.0782.7666.7465.8270.7885.67
Dense Latent2.720.790.7824.8331.1428.3451.6251.361.74
Dense Latent + TA3.473.262.0228.7629.3535.3652.0950.869.86
KV Ridge2.670.540.2623.5426.4525.6251.6251.611.67
KV Ridge + TA19.0937.4839.0774.7488.3168.5652.6465.6153.30
HeteroFold (Ours)33.7250.2943.1579.0090.3671.7663.2269.1663.46
Llama-3.1-8B $\rightarrow$ Ministral-3-14BTextMas45.5964.5652.0183.3292.6675.6361.5668.3394.69
Dense Latent3.131.331.0225.7933.1928.0049.8849.901.44
Dense Latent + TA6.714.884.4329.7238.5729.5850.5150.4118.65
KV Ridge1.240.250.8023.7822.7026.1651.3848.340.99
KV Ridge + TA11.4113.3315.4155.6675.0957.5354.9347.0847.99
HeteroFold (Ours)21.5839.0030.4372.0581.1465.9257.1466.4172.33

Table 2. Group decisions on HiddenBench. Initial accuracy precedes communication; the remaining metrics are measured after 15 rounds.

Communicationtabular[c]@r@Init.
Acc. (%)

Table 3. Component ablation for Ministral-3-14B$\rightarrow$Llama-3.1-8B. ARC-C, GSM8K and QuALITY report accuracy; Qasper, HotpotQA and LoCoMo report F1. The first three rows replace token alignment, depth aggregation, and cross-head mixing, respectively. All variants use the same calibration data as the main results; Calibration refers to our output-aware calibration.

MapARC-CGSM8KQasperHotpotQALoCoMoQuALITY
Same-index token pairing69.207.133.190.721.4724.59
Single-layer mapping82.0046.1722.3437.8730.4571.52
Head-local mapping71.8426.699.7117.3812.7554.79
TA + Ridge ($\ell_2$-regularized least squares)87.034.553.276.822.1347.41
+ key + value calibration89.1655.5721.6531.6631.9174.50
TA + Recolor90.1962.7726.5240.7140.8179.34
+ key calibration90.1962.3534.9948.7142.8578.41
+ value calibration90.5362.0225.9144.5040.5777.95
+ key + value calibration (HeteroFold)90.3663.4633.7250.2943.1579.00

Table 4. Transfer latency (in ms, $\downarrow$) and speedup over Native Prefill ($\times$, $\uparrow$) at batch size 2. All methods use FlashAttention-2. Latency sums synchronized stage timings across two NVIDIA H100 80GB GPUs connected by NVLink, including inter-GPU payload transfer for cache-transfer methods. Native Prefill processes the original text on the receiver. Parentheses report speedup over Native Prefill. Median of three runs after warm-up.

MethodLlama-3.1-8B $\rightarrow$ Ministral-3-14B 4,096 tokensLlama-3.1-8B $\rightarrow$ Ministral-3-14B 16,384 tokensLlama-3.1-8B $\rightarrow$ Ministral-3-14B 32,768 tokensQwen3-4B $\rightarrow$ Ministral-3-14B 4,096 tokensQwen3-4B $\rightarrow$ Ministral-3-14B 16,384 tokensQwen3-4B $\rightarrow$ Ministral-3-14B 32,768 tokens
Native Prefill443.1 (1.00$\times$)2091.2 (1.00$\times$)5167.4 (1.00$\times$)445.1 (1.00$\times$)2099.6 (1.00$\times$)5178.2 (1.00$\times$)
Dense Latent + TA155.4 (2.85$\times$)372.1 (5.62$\times$)705.7 (7.32$\times$)168.9 (2.64$\times$)389.3 (5.39$\times$)720.5 (7.19$\times$)
KV Ridge + TA135.6 (3.27$\times$)301.0 (6.95$\times$)569.9 (9.07$\times$)140.4 (3.17$\times$)316.5 (6.63$\times$)587.7 (8.81$\times$)
HeteroFold (Ours)118.3 (3.75$\times$)257.8 (8.11$\times$)481.3 (10.74$\times$)128.4 (3.47$\times$)290.1 (7.24$\times$)483.2 (10.72$\times$)

Table 5. Evaluation sets and metrics. Qasper and HotpotQA use LongBench subsets. LoCoMo excludes unanswerable category 5.

SuiteDatasetExamplesMetric
Short contextARC-Challenge1,172Accuracy
MMLU14,042Accuracy
WinoGrande1,267Accuracy
HellaSwag10,042Accuracy
GSM8K1,319Accuracy
Long contextQasper (LongBench)200Token F1
HotpotQA (LongBench)200Token F1
LoCoMo, categories 1–41,540Token F1
QuALITY, development2,086Accuracy
Multi-agentHiddenBench65Accuracy / invalid rate

Table 6. Sensitivity to calibration-set size with an equal Open-R1/HotpotQA mixture, correction rank 16, and three sender layers.

Training / held-out promptsARC-CHotpotQAQuALITY
200 / 5072.8726.4061.55
800 / 20075.6841.9763.57
1,600 / 400 (paper)77.0549.4462.90

Table 7. Sensitivity to correction rank with 1,600 training and 400 held-out prompts from the equal Open-R1/HotpotQA mixture and three sender layers. Rank 0 uses Recolor without output-aware calibration.

Correction rank $\rho_c$ARC-CHotpotQAQuALITY
0 (Recolor only)75.4330.4765.10
475.7747.6063.09
875.5146.7863.37
16 (paper)77.0549.4462.90
3276.7144.4564.57
6476.1143.0964.05

Table 8. Sensitivity to the sender-layer neighborhood with 1,600 training and 400 held-out prompts from the equal Open-R1/HotpotQA mixture and correction rank 16.

Sender-layer neighborhoodARC-CHotpotQAQuALITY
1 layer66.8922.5251.58
3 layers (paper)77.0549.4462.90
5 layers76.9648.4464.43

Figures in this post are taken from the original arXiv:2609.32259 (CC BY 4.0). Only size and format were changed.

License

Author: Jaehun Ryu

Link: https://jaehun.me/en/posts/paper-2609-32259/

License: CC BY 4.0

This work is licensed under the Creative Commons Attribution 4.0 International License. You are free to use it for any purpose, including commercial use, as long as you provide proper attribution.

Comments