Paper

Mechanics of Long-Context Hybrid Models: From Hybrid Attention to Hybrid Position

TL;DR — These days open-source LLMs are rapidly shifting from models that “use full attention only” to hybrid models that “mix full attention with Sliding-Window Attention (SWA) or Linear Attention (LA).” This paper mechanistically dissects “why” hybrid models work and “how” they can be designed better. The key conclusion is that the true identity of a hybrid model is in fact not hybrid attention but hybrid position (NoPE + position-biased attention). It identifies a Seesaw Effect — SWA hybrids are strong at length extrapolation but flip on context extension — and a No-Free-Lunch Effect — LA hybrids are strong within the training length but weak at extrapolation. As a solution it proposes Sliding-Window Linear Attention (SWLA), which extrapolates 4k→64k, i.e. 16× without training, while maintaining 100% accuracy on NIAH-SK1.


Key Idea

This is not a paper that proposes a single new model, but one that empirically organizes the mechanics of long-context hybrid LLMs. The authors’ reason for choosing the name “Mechanics” is clear: whereas prior work stopped at comparing which model is better, this paper places on a coordinate system how the different attention modules inside a model cooperate and how they affect length extrapolation and context extension (source: §1).

ItemValue
Model size376M / 776M / 1B (main) + 3B (validation) (source: §3.1, Tab.2)
Pretraining4k context × 50B tokens (source: §3.1)
Long-context continued training32k context × 5B tokens (source: §3.1)
SWA default window128, rotary base 10000 (source: §3.1)
Hybrid ratio3:1 (position-biased : NoPE) default (source: §3.1)
LA variantsGLA (Gated Linear Attention), GDN (Gated DeltaNet) (source: §3.1)
TokenizerLlama 3 family, vocab 128256 (source: Tab.2, App.A)
EvaluationPG19 (PPL), LongPPL, RULER, BABILong, LongBench (source: §3.1)
Hardware8× H200 (16× H200 for 3B) (source: §6.3, App.A)

The findings that run through the whole paper are condensed into six “Takeaways.” Extracting only the skeleton (source: §1, Takeaway List):

  • T1 · From Hybrid Attention to Hybrid Position — Hybrid models improve long-context performance and extrapolation through the combination of NoPE and other position biases.
  • T2 · Seesaw Effect — The long-context advantage SWA-NoPE had over LA-NoPE flips after long-context continued training.
  • T3 · No-Free-Lunch Effect — LA-NoPE is weaker than SWA-NoPE at direct extrapolation.
  • T4 · Tidal Effect — A small number of high-hit-rate NoPE handles “coarse global information gathering,” while many low-entropy position-biased attentions handle “noise removal.”
  • T5 · Short-Window Weariness / Long-Window Laziness — A short window is advantageous for extrapolation, a long window for extension.
  • T6 · Matthew Effect — Applying a sliding window to position-biased attention and strengthening global gathering for NoPE greatly improves extrapolation.

Background: The Problem They Solved

Architectural design in mainstream open-source LLMs is moving from RoPE-based full-attention-only models to hybrid models (source: §1, §2.1). Qwen-3.5, Kimi-K3, and OLMo-Hybrid mix full softmax attention (GQA/MLA) with linear attention (GDN/KDA), while Gemma4, Inkling, and Mimo-v2 mix full attention with sliding-window softmax attention (source: §2.1). Such models have good compute/memory efficiency and downstream performance that is similar or better (source: §2.1).

The research gap the authors point out is this. Prior work approached hybrid models from (1) downstream task evaluation, (2) scaling curves, and (3) system efficiency, but mechanistic analysis of “which module plays which role and why” is missing (source: §1, §2.2). Wang et al. reported that a 3:1 ratio is appropriate, and Wu et al. noted that retrieval is handled by full attention while SWA/LA affects the learning trajectory (source: §2.2), but they could not explain how modules with different positional inductive biases cooperate during extrapolation and extension.

Moreover, the million-token-scale context demands of the agent era place strong constraints on architecture: full attention alone is prohibitively expensive, and a single efficient attention alone is unreliable (source: §1). These three questions — “why hybrid” × “why long-context” × “why mechanistic” — are the starting point of this paper.


New Approach: Mechanics of Long-Context Hybrid Models

The authors run a closed-loop empirical study that trains, evaluates, and analyzes four model families under an identical protocol: (1) RoPE full attention, (2) RoPE-NoPE hybrid, (3) SWA hybrid, (4) GLA/GDN (LA) hybrid — each in two modes, layer-wise (LH) and head-wise (HH) (source: §1, §3.1). The structure is to infer causes from observations and verify the inferences through new improvements.

Methodologically, the two key inventions are two analysis tools (source: §4.2, §6.1):

  1. Entropy–Hit-Rate diagram. A 2D scatter plot that plots each attention head as a point. The x-axis is the attention entropy (normalized to 0–1) at 4k, and the y-axis is the top-64 hit rate. Retrieval heads and streaming heads are distinguished by color (source: §4.2).
  2. Extended attention entropy. LA has no softmax distribution, so entropy cannot be measured directly. The authors reinterpret the attention weight $\alpha_{t,s}$ as “the norm of the Jacobian of the output $o_t$ differentiated with respect to the value $v_s$” (source: §6.1):
$$ \alpha_{t,s} = \mathrm{softmax}\!\left(\frac{q_t k_s^{\top}}{\sqrt{d_k}}\right), \qquad o_t = \sum_{s=0}^{t} \alpha_{t,s} v_s, \qquad \left\lVert \frac{\partial o_t}{\partial v_s} \right\rVert_{2} = \alpha_{t,s} $$

In softmax attention, $\alpha_{t,s}$ is already normalized, so the result stays the same, and in LA it leaves the original computation untouched while allowing SWA, GLA, GDN, and RoPE to be placed on a single diagram in the common language of “entropy” (source: §6.1).


How It Works: A Concrete Example Walkthrough

Example 1 — NoPE and RoPE Do Different “Jobs”

Suppose a model trained at 4k must find a single line hidden somewhere in a 64k document: “the password is 8842” (RULER’s NIAH-SK1 task, source: §4.2). Here (source: §4.2, Fig.12):

  • NoPE attention has a high top-64 hit rate, a high proportion of retrieval heads, and high entropy. That is, it is a global detector that coarsely pinpoints “roughly where the needle is.” However, high entropy also means that attention is spread widely, so if it cannot pinpoint precisely, it drags in noise instead.
  • RoPE (and SWA, GLA, GDN) attention has a low hit rate, a high proportion of streaming heads, and low entropy. It handles narrow, clean local reading. Thanks to positional bias it cannot see far, but in the region it does cover there is little noise.

This relationship is precisely the Tidal Effect (T4) (source: §4.2, Fig.13). In the entropy–hit-rate diagram, position-biased attention occupies a “sector spreading from the lower-left corner,” while NoPE occupies a “strip running from upper-center to lower-right.” Raising the NoPE ratio from 3:1 → 1:1 → 1:3 makes the boundary between the two regions shift (tidal lane), and when the NoPE ratio becomes excessive, performance drops instead (source: §4.1). In the noise experiment too, injecting Gaussian noise into position-biased heads degrades performance more than injecting it into NoPE — paradoxical evidence that NoPE alone cannot select information precisely (source: §4.1, Fig.11).

Example 2 — Sliding-Window Linear Attention (SWLA), Step by Step

The reason LA hybrids are weak at extrapolation is that the linear attention gate induces a data-dependent positional bias (source: §1, §6.1). Taking GLA as an example, behind the interaction of the query $q_t$ and key $k_s$ there is a product of gates (source: §6.2):

$$ q_t k_s^{\top} \prod_{i=s+1}^{t} \mathrm{Diag}(\alpha_i) $$

This product carries the relative position signal, but as the context grows longer, the number of accumulated factors grows without bound, pushing the position signal outside the training range (4k). SWLA confines this relative position signal inside a window. Concretely, with window 4096 and stride 1024 (source: §6.2):

  1. Use the original GLA/GDN operator (FLA kernel) as is to compute chunks $[0,4095]$, $[1024,5119]$, $[2048,6143]$, … in parallel.
  2. From each chunk, keep only the valid interval — $[0,4095]$, $[4096,5119]$, $[5120,6143]$, … — and crop the rest.
  3. Concatenate the cropped pieces back together.

This approach does not depend on the internal implementation of a particular LA variant, making it “compatible with any LA variant” (source: §6.2). Adding to this a log-scale correction for NoPE attention (increasing the attention logit scale logarithmically to suppress entropy growth, source: §3.2) yields EME (Extrapolation based on Matthew Effect) (source: §6.2). The gist is simple: confine the relative position of position-biased attention, and strengthen the global gathering of NoPE.


Performance Validation: Main Results

Immediately after short-context training (4k). Compared to RoPE-only models, hybrid models show lower loss and PPL within the training length, and SWA/GLA/GDN-NoPE hybrids show long-context performance that surpasses RoPE or RoPE-NoPE hybrids (source: §3.2, Fig.1–4). Notable is that a RoPE-NoPE hybrid alone becomes a training-free extrapolation of a model trained with full attention (T1) (source: §3.2, Fig.5).

Seesaw Effect (T2). After long-context continued training (32k), the ranking flips. SWA-NoPE, previously the strongest at extrapolation, is overtaken by LA-NoPE after context extension, and sometimes it even fails to exceed the extrapolation performance of log-scale NoPE — this is especially pronounced in layer-wise (source: §3.3, Fig.8). The cause is the Short-Context Learning Trap (source: §5.1): SWA-NoPE is already good at extrapolation, so during long-context continued training the loss “barely decreases at long positions and drops abnormally sharply only at short positions.” Looking at the rate of parameter change, in SWA/NoPE hybrids the rate of change of SWA’s $W_Q$·$W_K$ is far larger than NoPE’s — but since the SWA window is fixed at 128, in the end it only learns narrow local branches (source: §5.1, Fig.15). The authors resolved this by enlarging the window from 256→4096 and adding a LongCE loss, and window 2048 was optimal (source: §5.2, Fig.16). This is summarized as T5: the need for “stage-wise positional bias tuning” — a short window is advantageous for short-context learning (extrapolation), a long window for long-context learning (extension) (source: §5.2).

No-Free-Lunch Effect (T3) and SWLA (T6). LA-NoPE is strong at fitting within the training length but weak at direct extrapolation (source: §3.3). Applying SWLA greatly improves the extrapolation of GLA/GDN-NoPE, and the log-scale correction only then takes effect (source: §6.2, Fig.23). The decisive numbers are the NIAH task for the 776M model (source: §6.2, Tab.1):

776MSK1 @64kSK2 @64kSK3 @64k
GLA-NoPE-LH (Short)0.00.00.0
+ SWLA~97~0~3
+ SWLA & Log (EME)~100~59~64

That is, under 16× extrapolation (4k→64k), SK1 maintains 100%, but as we move to SK2/SK3 with multiple needles, difficulty rises sharply and performance drops (source: Tab.1). GDN-NoPE-LH was already strong at 99% on SK1 from the baseline, and with EME it holds up to 90% on SK2 and 53% on SK3 (source: Tab.1). The four models — RoPE-NoPE, GLA-NoPE, GDN-NoPE, and SWA-NoPE — that extrapolate with EME show broadly similar levels (source: §6.2, Fig.25). The LongBench validation in the appendix also reproduces the Seesaw Effect and the Matthew Effect (source: App.A.1).

Efficiency. Hybrid models show a clear advantage over full attention in TTFT (Time to First Token), TPOT (Time per Output Token), peak memory, and cache size as the model and context grow (source: §6.3, Fig.26). Training throughput (TGS, tokens/GPU/s) is also larger for hybrids at 1B·32k (source: §6.3, Fig.27). However, GLA/GDN store the recurrent state in FP32, so the cache is at a level similar to SWA (w=128, FP16), and head-wise SWA does not split the QKV but concatenates it, so intermediate memory is large and peak memory equals that of full attention (source: §6.3).


Our Perspective: Strengths, Limitations, and Why This Research Matters

Strengths. The greatest value of this paper is that it moves design principles onto a quantified coordinate system. It validates the existing vague intuition that “NoPE searches, RoPE removes noise” on the common coordinate of the entropy–hit-rate diagram, and adds to it a dynamics in which “the boundary shifts with the hybrid ratio” (Tidal Effect) (source: §4.2). In particular, extended entropy, which makes even LA analyzable, is a clean contribution generally applicable to attention without softmax (source: §6.1). And SWLA is implemented as a chunk-wise approximation without modifying kernels, giving it the practicality of being applicable to any LA variant (source: §6.2).

Limitations. To be fair, though, a few points must be noted.

  1. Small scale. 376M–3B is tens of times smaller than production hybrid models (Qwen-3.5, Kimi-K3 at tens to hundreds of B). Generalization of the findings is only “at this scale.” Although the authors themselves do not admit this limitation, validation on scaling curves is lacking.
  2. The trap of “100% SK1.” The headline number of 100% is confined to the easiest single-needle task (SK1). On SK2/SK3 it plunges to 59%·64% (source: Tab.1). This reminds us that the phrase “16× extrapolation” refers to the extrapolation of a specific retrieval task, not overall long-context ability.
  3. Confined to pretraining. As the authors explicitly state, the analysis scope is concentrated on the pretraining stage, and post-training and multimodality (especially vision-language) are unverified (source: Limitations). The architectural scope also omits other hybrid families such as sparse attention (DSA) and compressed attention (DeepSeek-v4) (source: Limitations).
  4. The inherent instability of NoPE itself. The conclusion that “NoPE steadily gathers global information but cannot localize precisely, so cooperation is needed” (the full version of T1, source: §4.2) means it is sensitive to the NoPE ratio and placement, which leads to hyperparameter search burden.

Why it matters nonetheless. Hybrid models have already become a default choice across the industry — Jamba, GPT-OSS, MiniMax-01, Qwen, Kimi, Nemotron-H, and others (source: §2.1). In this trend, providing a shared language and coordinate system for “why it works” is work that raises architectural design one step from the rules of thumb of individual papers to predictable engineering. The Matthew Effect principle, “confine positional bias, broaden NoPE,” is itself an immediately usable design principle.


What’s Next?: The Road Ahead

The authors suggest three directions (source: Limitations): (1) expanding validation to post-training and multimodality, (2) applying the same mechanistic lens to other hybrid families such as sparse attention and compressed attention, and (3) analyzing detailed design choices such as sink bias, partial RoPE, gate variants, and short-convolution ablations.

Adding to this the next steps that would reasonably follow:

  • Scale validation. How well the Seesaw/Tidal/Matthew effects hold at 7B–70B. In particular, whether the cause of the Seesaw Effect (loss concentrating on short positions) worsens or disappears with scale is practically decisive.
  • Theorizing SWLA. Connecting the intuition “confine the relative position signal to a window” with position designs such as FoX, DeltaFormer, and PaTH, and deriving a quantitative relationship between window size and rotary base·gate decay rate, would allow the stage-wise tuning (T5) to be automated.
  • An integrated retrieval–tracking metric. The gap between SK1’s 100% and the plunge on SK2/SK3 shows that “single-needle retrieval ≠ long-context understanding.” An integrated metric that looks at multi-needle and state-tracking together is needed.

Ultimately, the promise of this series (Part 1.1 → follow-ups) is to create a “design-principle manual” for hybrid models. As its first chapter, this paper has opened a sufficiently convincing door.


Tables from the paper

Tables converted mechanically from the arXiv e-print LaTeX source. The numbers are the paper’s own and did not pass through a model.

Table 1. Extrapolation comparison of the 776M GLA/GDN-NoPE hybrid model on the NIAH-SK1, SK2, and SK3 tasks. Applying a sliding-window linear attention and a log scale for NoPE attention achieves the best length extrapolation. tab-niah-matthew-la

NIAH-SK1 4kNIAH-SK1 8kNIAH-SK1 16kNIAH-SK1 32kNIAH-SK1 64kNIAH-SK2 4kNIAH-SK2 8kNIAH-SK2 16kNIAH-SK2 32kNIAH-SK2 64kNIAH-SK3 4kNIAH-SK3 8kNIAH-SK3 16kNIAH-SK3 32kNIAH-SK3 64k
776M Short
GLA-NoPE-LH99.095.080.00.00.099.069.00.00.00.090.088.00.00.00.0
+ Log99.099.098.00.00.099.068.01.00.00.090.091.00.00.00.0
+ SWLA99.0100.0100.097.097.099.074.05.00.00.090.056.021.08.03.0
+ SWLA & Log (EME)99.0100.0100.096.0100.099.092.087.072.059.090.079.075.076.064.0
\hdashline GLA-NoPE-HH100.0100.068.082.00.0100.071.00.00.00.065.045.00.00.00.0
+ Log100.0100.098.095.072.0100.068.00.00.00.065.049.00.00.00.0
+ SWLA100.0100.0100.092.075.0100.0100.082.032.01.065.062.044.09.01.0
+ SWLA & Log (EME)100.0100.0100.0100.0100.0100.0100.097.066.031.065.068.065.062.034.0
\hdashline GDN-NoPE-LH100.0100.0100.0100.099.099.098.02.00.00.062.071.05.00.00.0
+ Log100.0100.0100.0100.0100.099.098.09.01.00.062.070.09.00.00.0
+ SWLA100.0100.0100.0100.099.099.097.096.092.053.062.064.061.056.031.0
+ SWLA & Log (EME)100.0100.0100.0100.0100.099.097.099.099.090.062.067.064.066.053.0
\hdashline GDN-NoPE-HH100.0100.0100.099.092.0100.095.02.00.00.094.038.00.00.00.0
+ Log100.0100.0100.0100.0100.0100.093.03.00.00.094.043.00.00.00.0
+ SWLA100.099.0100.095.092.0100.099.064.04.00.094.068.031.05.02.0
+ SWLA & Log (EME)100.0100.0100.0100.0100.0100.0100.097.072.035.094.071.065.047.016.0

Table 2. The hyperparameters of different model sizes.tab-config

376M776M1B3B
Hidden Size1024153620483072
Intermediate Size35845376716810752
Num Layer8121624
Num Attn Head8121624
Num KV Head8121624
Vocab Size128256128256128256128256

Table 3. Short-context performance of layer-wise hybrid models after short-context pretraining.tab-short-main-lh-short

TQALMDPIQAHellaWinoARC-eARC-cGPQASIQAOBQASGMMLUAvg.
376M Short
RoPE36.840.265.933.852.239.726.426.339.227.442.124.637.9
RoPE-NoPE-LH36.441.765.933.851.139.227.528.840.325.845.226.938.5
\hdashline SWA-RoPE-LH35.842.665.833.452.337.025.424.839.927.644.825.937.9
SWA-NoPE-LH36.442.066.433.450.237.224.125.839.721.643.625.637.2
\hdashline GLA-RoPE-LH37.042.766.433.852.338.825.124.239.026.644.726.338.1
GLA-NoPE-LH36.044.265.934.152.439.723.423.738.927.444.924.938.0
\hdashline GDN-RoPE-LH36.145.366.534.452.938.824.126.840.627.643.625.038.5
GDN-NoPE-LH36.044.965.134.450.537.725.428.839.827.243.526.438.3
776M Short
RoPE35.851.368.442.052.745.927.125.340.626.643.024.640.3
RoPE-NoPE-LH36.450.968.841.552.041.326.427.340.424.643.225.139.8
\hdashline SWA-RoPE-LH36.749.869.942.251.644.430.528.340.727.445.225.241.0
SWA-NoPE-LH36.350.769.841.753.942.328.824.240.926.645.726.440.6
\hdashline GLA-RoPE-LH36.852.169.242.253.745.229.526.841.027.843.525.641.1
GLA-NoPE-LH37.352.669.842.552.641.331.223.740.124.243.824.940.3
\hdashline GDN-RoPE-LH37.753.968.942.253.842.232.219.240.425.644.625.940.5
GDN-NoPE-LH37.352.669.442.954.642.029.827.840.826.445.426.541.3
1B Short
RoPE38.756.570.846.855.347.629.529.341.322.044.526.842.4
RoPE-NoPE-LH36.857.471.648.655.348.928.821.240.723.643.226.141.8
\hdashline SWA-RoPE-LH37.457.970.247.653.446.431.524.241.723.444.326.242.0
SWA-NoPE-LH37.957.171.346.852.847.331.923.741.125.845.326.642.3
\hdashline GLA-RoPE-LH37.457.471.048.656.347.131.525.341.623.243.126.442.4
GLA-NoPE-LH36.759.171.948.455.246.431.923.241.025.644.525.142.4
\hdashline GDN-RoPE-LH37.959.271.847.956.647.629.821.741.822.246.625.242.4
GDN-NoPE-LH37.757.771.247.255.046.430.928.341.524.647.423.442.6
3B Short
RoPE36.662.973.355.258.051.733.625.841.727.645.425.644.8
RoPE-NoPE-LH36.063.573.656.158.052.629.821.242.627.646.526.744.5
\hdashline SWA-RoPE-LH36.662.973.655.358.654.731.925.342.322.047.826.044.7
SWA-NoPE-LH37.362.273.555.757.353.632.521.742.023.645.026.544.3
\hdashline GLA-RoPE-LH37.662.373.152.856.949.932.525.842.025.446.725.644.2
GLA-NoPE-LH36.864.273.754.557.450.333.229.342.026.247.725.045.0
\hdashline GDN-RoPE-LH37.065.473.754.860.653.832.222.741.626.248.626.445.3
GDN-NoPE-LH36.164.274.256.157.755.730.921.242.427.046.826.544.9

Table 4. Short-context performance of head-wise hybrid models after short-context pretraining.tab-short-main-hh-short

TQALMDPIQAHellaWinoARC-eARC-cGPQASIQAOBQASGMMLUAvg.
376M Short
RoPE36.840.265.933.852.239.726.426.339.227.442.124.637.9
RoPE-NoPE-HH37.140.166.033.350.440.024.827.839.927.843.325.838.0
\hdashline SWA-RoPE-HH35.141.365.433.651.339.723.722.738.828.045.026.337.6
SWA-NoPE-HH35.841.765.533.851.739.026.427.838.723.442.225.237.6
\hdashline GLA-RoPE-HH36.046.066.035.052.440.224.825.839.123.643.525.638.2
GLA-NoPE-HH36.045.265.734.651.537.725.827.839.627.844.726.238.5
\hdashline GDN-RoPE-HH36.844.966.334.751.338.324.826.840.628.044.026.738.6
GDN-NoPE-HH35.844.267.134.752.639.723.430.839.126.446.125.438.8
776M Short
RoPE35.851.368.442.052.745.927.125.340.626.643.024.640.3
RoPE-NoPE-HH35.251.668.942.454.342.227.124.840.121.243.625.339.7
\hdashline SWA-RoPE-HH37.349.767.941.552.040.726.424.239.618.842.826.138.9
SWA-NoPE-HH35.850.869.742.353.442.226.825.340.024.844.326.440.1
\hdashline GLA-RoPE-HH38.351.970.242.452.343.729.826.340.931.642.925.941.4
GLA-NoPE-HH36.053.070.142.952.345.528.124.839.723.244.025.340.4
\hdashline GDN-RoPE-HH36.452.070.743.553.046.430.224.240.721.844.225.540.7
GDN-NoPE-HH37.654.068.843.255.245.729.521.240.327.043.324.240.8
1B Short
RoPE38.756.570.846.855.347.629.529.341.322.044.526.842.4
RoPE-NoPE-HH36.657.070.047.655.046.430.225.841.227.845.925.142.4
\hdashline SWA-RoPE-HH38.756.370.647.255.346.632.227.340.924.644.125.942.5
SWA-NoPE-HH36.156.171.147.955.046.630.922.241.026.845.126.342.1
\hdashline GLA-RoPE-HH36.158.870.649.654.548.931.920.241.726.643.326.342.4
GLA-NoPE-HH36.759.171.048.755.847.832.224.841.326.644.624.942.8
\hdashline GDN-RoPE-HH36.157.972.049.654.746.029.821.741.827.246.726.642.5
GDN-NoPE-HH35.258.870.149.353.647.431.520.741.819.647.825.741.8
3B Short
RoPE36.662.973.355.258.051.733.625.841.727.645.425.644.8
RoPE-NoPE-HH36.363.573.156.057.549.932.528.342.023.846.924.444.5
\hdashline SWA-RoPE-HH37.163.273.756.358.451.235.326.842.425.846.427.645.3
SWA-NoPE-HH36.462.473.254.458.551.531.521.742.322.445.925.743.8
\hdashline GLA-RoPE-HH36.663.373.756.259.051.031.924.243.120.846.427.444.5
GLA-NoPE-HH35.463.974.256.358.350.631.224.242.728.848.625.945.0
\hdashline GDN-RoPE-HH37.764.174.557.559.851.234.927.343.524.045.025.645.4
GDN-NoPE-HH36.863.973.754.760.350.837.329.342.625.848.124.845.7

Table 5. Short-context performance of layer-wise hybrid models after long-context continual pretraining.tab-short-main-lh-long

TQALMDPIQAHellaWinoARC-eARC-cGPQASIQAOBQASGMMLUAvg.
376M Long
RoPE35.235.864.331.251.236.725.824.838.528.043.624.236.6
RoPE-NoPE-LH36.038.764.231.752.537.628.825.839.228.242.924.737.5
\hdashline SWA-RoPE-LH35.439.364.231.052.439.225.823.738.328.045.325.237.3
+ DRoPE35.539.165.030.851.639.023.424.838.127.644.725.737.1
SWA-NoPE-LH35.839.664.331.449.036.725.425.337.626.444.425.036.7
\hdashline GLA-RoPE-LH37.141.664.432.052.237.024.125.339.127.244.926.237.6
+ DRoPE36.040.263.731.651.237.625.126.338.427.443.425.737.2
GLA-NoPE-LH37.641.664.431.751.538.623.425.838.027.844.324.637.4
\hdashline GDN-RoPE-LH36.043.763.832.252.839.023.425.838.727.644.224.837.7
+ DRoPE35.843.364.132.152.039.322.724.239.027.644.725.137.5
GDN-NoPE-LH36.441.964.532.151.237.625.424.239.028.044.024.937.4
776M Long
RoPE37.747.966.138.752.641.126.125.839.627.643.124.639.2
RoPE-NoPE-LH37.048.367.038.451.841.526.825.840.327.642.924.639.3
\hdashline SWA-RoPE-LH36.445.167.738.651.942.225.417.739.027.645.324.038.4
+ DRoPE36.345.767.838.752.041.825.821.239.026.444.923.938.6
SWA-NoPE-LH37.045.866.738.352.341.528.124.839.827.845.524.339.3
\hdashline GLA-RoPE-LH37.649.167.739.053.841.627.824.240.326.845.524.739.8
+ DRoPE37.349.768.038.754.642.025.824.240.127.245.025.239.8
GLA-NoPE-LH36.850.667.239.252.341.830.220.740.027.645.025.639.7
\hdashline GDN-RoPE-LH37.451.466.940.153.542.029.825.340.627.444.325.040.3
+ DRoPE36.650.665.939.252.541.328.824.240.227.044.124.239.5
GDN-NoPE-LH37.051.168.040.152.839.927.125.340.427.443.825.739.9
1B Long
RoPE38.754.469.243.853.844.127.525.840.327.045.427.841.5
RoPE-NoPE-LH37.954.069.344.654.745.728.124.240.025.844.826.641.3
\hdashline SWA-RoPE-LH37.054.769.644.754.548.028.524.840.727.644.627.341.8
+ DRoPE37.154.168.944.052.646.228.824.840.928.045.026.641.4
SWA-NoPE-LH39.053.469.244.251.943.930.523.241.527.845.427.241.4
\hdashline GLA-RoPE-LH37.755.269.245.855.546.429.824.840.526.043.926.741.8
+ DRoPE35.755.069.445.355.544.128.526.340.627.043.426.241.4
GLA-NoPE-LH37.455.869.845.554.244.429.226.840.627.843.323.541.5
\hdashline GDN-RoPE-LH38.056.370.245.856.346.428.822.239.827.244.426.341.8
+ DRoPE37.955.971.245.856.347.429.822.240.227.845.025.742.1
GDN-NoPE-LH37.355.869.245.155.844.829.825.840.827.045.125.741.9
3B Long
RoPE36.659.971.650.456.346.429.223.741.528.245.827.243.1
RoPE-NoPE-LH36.461.172.551.357.147.630.923.241.527.446.126.143.4
\hdashline SWA-RoPE-LH37.461.372.152.457.249.930.527.341.127.648.127.344.3
+ DRoPE37.660.671.852.456.150.632.528.841.425.448.026.744.3
SWA-NoPE-LH36.060.472.552.257.150.330.519.741.827.245.924.543.2
\hdashline GLA-RoPE-LH37.461.472.953.458.551.033.227.840.326.446.626.544.6
+ DRoPE37.161.472.954.256.851.731.225.840.726.246.523.944.0
GLA-NoPE-LH37.162.372.351.555.448.034.221.741.725.846.026.543.5
\hdashline GDN-RoPE-LH38.563.773.152.159.352.631.224.840.523.249.026.544.5
+ DRoPE37.663.373.051.960.952.431.224.840.425.446.526.244.5
GDN-NoPE-LH36.662.272.553.156.452.432.927.841.826.846.524.644.5

Table 6. Short-context performance of head-wise hybrid models after long-context continual pretraining.tab-short-main-hh-long

TQALMDPIQAHellaWinoARC-eARC-cGPQASIQAOBQASGMMLUAvg.
376M Long
RoPE35.235.864.331.251.236.725.824.838.528.043.624.236.6
RoPE-NoPE-HH37.936.463.631.350.238.527.126.337.828.044.126.137.3
\hdashline SWA-RoPE-HH35.237.663.331.451.037.022.424.837.928.045.125.136.6
+ DRoPE35.137.163.331.352.038.123.424.837.227.845.125.236.7
SWA-NoPE-HH35.536.963.831.551.737.226.125.837.927.843.823.536.8
\hdashline GLA-RoPE-HH37.043.065.432.451.539.224.425.838.224.643.925.437.6
+ DRoPE36.342.664.432.251.637.223.425.338.326.844.725.137.3
GLA-NoPE-HH36.342.463.332.651.938.125.124.239.127.444.725.437.5
\hdashline GDN-RoPE-HH34.941.563.932.952.338.823.424.839.028.043.125.737.4
+ DRoPE36.740.364.332.451.539.223.725.339.228.243.725.737.5
GDN-NoPE-HH36.139.164.632.651.638.125.124.238.327.643.626.137.2
776M Long
RoPE37.747.966.138.752.641.126.125.839.627.643.124.639.2
RoPE-NoPE-HH35.747.467.738.553.840.924.425.839.022.644.826.238.9
\hdashline SWA-RoPE-HH36.746.666.338.851.342.227.124.839.327.043.625.039.1
+ DRoPE37.445.866.438.252.640.727.126.339.526.644.324.439.1
SWA-NoPE-HH35.846.367.239.152.843.024.125.340.327.643.024.439.1
\hdashline GLA-RoPE-HH38.349.168.239.453.442.325.425.840.328.044.025.440.0
+ DRoPE38.048.468.739.552.941.826.125.340.828.045.326.840.1
GLA-NoPE-HH37.050.368.639.852.642.325.826.339.726.644.026.039.9
\hdashline GDN-RoPE-HH36.448.469.240.153.043.627.526.340.727.445.724.540.2
+ DRoPE35.748.369.639.953.341.330.526.840.328.245.225.440.4
GDN-NoPE-HH38.951.667.940.255.044.827.819.739.227.645.524.240.2
1B Long
RoPE38.754.469.243.853.844.127.525.840.327.045.427.841.5
RoPE-NoPE-HH36.853.669.344.854.043.728.524.239.728.044.825.141.1
\hdashline SWA-RoPE-HH38.551.869.845.053.745.728.825.841.627.244.324.741.4
+ DRoPE38.752.170.144.452.943.927.526.341.728.244.824.741.3
SWA-NoPE-HH37.954.669.244.755.243.229.219.739.825.044.527.240.8
\hdashline GLA-RoPE-HH37.454.269.545.252.647.128.120.741.128.044.125.241.1
+ DRoPE38.054.269.945.653.245.727.123.740.527.644.324.841.2
GLA-NoPE-HH36.856.270.245.553.947.830.924.241.027.844.326.442.1
\hdashline GDN-RoPE-HH36.755.270.745.954.144.330.225.340.927.445.725.841.8
+ DRoPE37.055.370.145.353.343.028.126.341.028.245.626.641.7
GDN-NoPE-HH37.656.869.445.254.145.230.220.740.827.446.126.041.6
3B Long
RoPE36.659.971.650.456.346.429.223.741.528.245.827.243.1
RoPE-NoPE-HH35.861.172.152.456.949.730.523.240.522.844.125.542.9
\hdashline SWA-RoPE-HH36.661.172.051.857.650.632.521.741.223.046.028.643.6
+ DRoPE37.060.672.052.456.849.930.225.841.324.045.027.543.5
SWA-NoPE-HH36.459.972.551.958.250.629.524.841.825.246.025.543.5
\hdashline GLA-RoPE-HH36.161.172.953.456.851.331.923.741.526.444.726.743.9
+ DRoPE37.360.372.953.456.850.130.924.241.024.445.526.343.6
GLA-NoPE-HH35.861.673.453.157.749.930.528.342.325.845.526.644.2
\hdashline GDN-RoPE-HH37.662.772.053.858.250.333.228.842.126.245.826.544.8
+ DRoPE37.462.073.153.157.349.033.927.842.023.446.624.344.2
GDN-NoPE-HH36.862.472.352.560.651.035.325.841.826.045.925.744.7

Table 7. Long-context performance of 376M and 776M layer-wise hybrid models after short-context pretraining.tab-ruler-lh-short-small

RULER 4kRULER 8kRULER 16kRULER 32kRULER 64kBABILong 0kBABILong 2kBABILong 4kBABILong 8kBABILong 16kBABILong 32kBABILong 64kAverage RU.Average BAAverage $\leq$4kAverage >4kAverage All
376M Short
RoPE25.80.00.00.00.021.717.88.00.40.00.00.05.26.818.30.16.1
+ NTK26.59.90.60.10.021.817.87.88.92.20.60.07.48.418.52.88.0
+ NTK + Log26.514.01.90.90.521.817.88.07.47.92.20.68.89.418.54.49.1
RoPE-NoPE-LH29.90.10.00.00.032.919.515.70.60.20.00.06.09.824.50.18.2
+ NTK30.612.11.30.70.433.019.515.923.24.90.70.09.013.924.75.411.9
+ NTK + Log30.618.72.20.60.333.019.415.928.418.45.70.910.517.424.79.414.5
SWA-4k-NoPE-LH29.918.06.62.82.233.119.515.724.418.713.77.911.919.024.611.816.0
+ Log (EME)29.918.411.56.63.333.019.515.726.826.725.924.714.024.624.518.020.2
\hdashline SWA-RoPE-LH20.86.00.50.20.232.126.513.09.45.84.43.95.513.623.13.810.2
SWA-NoPE-LH29.123.519.311.09.231.027.223.920.717.112.912.618.420.827.815.819.8
+ Log29.124.222.524.919.431.427.223.923.422.019.319.024.023.727.921.823.9
\hdashline GLA-RoPE-LH24.23.80.70.20.038.122.613.59.48.86.75.15.814.924.64.311.1
GLA-NoPE-LH28.617.02.01.30.639.037.033.825.511.32.62.89.921.734.67.916.8
+ Log28.619.02.61.30.739.037.033.920.69.73.03.110.420.934.67.516.5
+ wsz=4k28.523.116.910.24.538.837.033.828.015.613.012.916.725.634.515.521.9
+ EME28.623.726.526.019.938.837.033.931.427.322.919.224.930.134.624.627.9
\hdashline GDN-RoPE-LH21.22.00.90.50.036.831.916.613.54.21.40.34.915.026.62.910.8
GDN-NoPE-LH22.920.410.46.03.631.834.030.327.011.81.70.112.719.529.710.116.7
+ Log22.921.214.67.96.931.934.030.328.716.33.80.214.720.729.812.518.2
+ wsz=4k22.921.518.115.57.632.034.030.328.228.721.716.917.127.429.819.823.1
+ EME22.921.922.720.813.631.934.030.329.029.528.426.220.429.929.824.025.9
776M Short
RoPE34.90.50.00.00.043.939.826.51.60.70.10.07.116.136.30.412.3
+ NTK36.828.62.50.31.843.939.728.418.45.54.82.414.020.437.28.017.8
+ NTK + Log36.930.412.32.81.244.139.828.222.710.63.02.916.721.637.210.719.6
RoPE-NoPE-LH32.10.30.00.00.041.134.730.91.90.32.30.06.515.934.70.612.0
+ NTK32.317.80.50.10.141.234.631.624.55.91.90.510.220.034.96.415.9
+ NTK + Log32.326.15.01.50.341.034.731.329.320.76.94.813.024.134.811.819.5
SWA-4k-NoPE-LH32.222.715.53.01.441.034.730.831.027.221.918.515.029.334.717.723.3
+ Log (EME)32.224.29.77.02.541.134.630.830.629.325.623.315.130.834.719.024.2
\hdashline SWA-RoPE-LH22.911.43.21.10.846.331.420.516.812.29.37.57.920.630.37.815.3
SWA-NoPE-LH34.929.423.211.79.935.333.630.528.223.420.516.821.826.933.620.424.8
+ Log34.832.531.726.625.235.333.730.530.626.924.724.230.229.433.627.829.7
\hdashline GLA-RoPE-LH31.66.71.51.10.546.334.924.719.513.711.03.08.321.934.47.116.2
GLA-NoPE-LH35.028.710.50.50.940.436.333.428.714.92.52.615.122.736.311.219.5
+ Log35.029.211.50.70.640.336.333.531.217.62.42.515.423.436.312.020.1
+ wsz=4k35.226.915.211.210.940.436.333.426.719.014.812.819.926.236.317.223.6
+ EME35.132.130.526.122.340.336.333.431.328.424.320.829.230.736.327.030.1
\hdashline GDN-RoPE-LH31.29.12.01.51.142.132.722.916.113.210.22.19.019.932.26.915.3
GDN-NoPE-LH37.333.415.511.68.838.232.032.122.212.49.53.921.421.534.914.721.4
+ Log37.333.716.711.18.638.332.032.021.910.18.06.721.521.334.914.621.4
+ wsz=4k37.334.434.330.421.238.232.032.027.221.917.415.631.526.334.925.328.5
+ EME37.234.734.633.028.338.332.032.129.326.825.722.033.629.534.929.331.2

Table 8. Long-context performance of 1B and 3B layer-wise hybrid models after short-context pretraining.tab-ruler-lh-short-large

RULER 4kRULER 8kRULER 16kRULER 32kRULER 64kBABILong 0kBABILong 2kBABILong 4kBABILong 8kBABILong 16kBABILong 32kBABILong 64kAverage RU.Average BAAverage $\leq$4kAverage >4kAverage All
1B Short
RoPE37.31.50.00.00.047.441.728.81.70.30.20.17.817.238.80.513.3
+ NTK39.329.76.90.50.747.541.628.821.59.07.02.015.422.539.39.719.5
+ NTK + Log39.333.418.78.65.847.441.728.824.720.013.49.521.226.539.316.824.3
RoPE-NoPE-LH46.90.10.20.00.065.251.948.12.50.10.00.09.424.053.00.417.9
+ NTK45.835.96.90.30.265.151.747.039.313.34.20.217.831.552.412.525.8
+ NTK + Log45.837.523.75.21.465.351.947.045.441.930.89.922.741.752.524.533.8
SWA-4k-NoPE-LH46.940.439.534.122.265.351.948.139.031.824.920.036.640.153.131.538.7
+ Log (EME)46.941.141.342.038.065.351.948.140.538.534.328.741.943.953.138.143.1
\hdashline SWA-RoPE-LH34.814.23.31.10.963.342.828.121.39.08.39.610.926.142.28.519.7
SWA-NoPE-LH44.838.129.513.09.953.940.737.031.429.821.917.127.133.144.123.830.6
+ Log44.842.643.040.935.153.840.737.034.534.029.027.041.336.644.135.838.5
\hdashline GLA-RoPE-LH37.112.30.50.50.247.747.531.619.45.86.24.210.123.241.06.117.8
GLA-NoPE-LH47.040.313.53.21.542.747.245.438.920.39.88.721.130.445.617.026.6
+ Log47.141.220.74.32.342.847.345.440.720.111.510.823.131.245.718.927.8
+ wsz=4k47.040.237.235.316.842.647.345.438.330.520.917.335.334.645.629.634.9
+ EME47.042.141.542.537.542.847.345.440.334.931.929.342.138.845.637.540.2
\hdashline GDN-RoPE-LH41.316.03.81.60.761.350.232.820.713.03.85.712.726.846.48.220.9
GDN-NoPE-LH44.727.513.811.69.638.636.837.631.729.215.614.721.429.239.419.225.9
+ Log44.631.114.811.19.438.636.837.634.029.916.914.622.229.839.420.226.6
+ wsz=4k44.736.036.627.818.638.636.837.634.927.921.216.132.730.439.427.431.4
+ EME44.738.841.438.233.538.636.837.638.132.128.826.439.334.139.434.736.2
3B Short
RoPE42.41.60.10.00.056.855.740.45.90.70.00.08.822.848.81.017.0
+ NTK43.734.615.21.10.756.656.039.828.915.16.40.419.129.049.012.824.9
+ NTK + Log43.837.032.119.66.557.155.939.530.726.120.414.727.834.949.123.431.9
RoPE-NoPE-LH56.620.50.10.00.067.458.752.029.03.22.40.415.530.458.77.024.2
+ NTK57.140.019.12.40.867.358.750.742.625.07.83.123.936.558.517.631.2
+ NTK + Log57.045.734.313.35.866.958.650.745.941.432.816.731.244.758.329.539.1
SWA-4k-NoPE-LH56.540.236.633.327.167.558.752.245.841.130.420.438.745.258.734.442.5
+ Log (EME)56.940.336.231.925.767.458.652.144.743.238.532.738.248.258.836.744.0
\hdashline SWA-NoPE-LH51.846.944.635.522.964.356.548.941.736.631.022.440.343.155.435.241.9
+ Log51.948.948.845.944.063.956.448.844.843.137.232.547.946.755.243.247.2
\hdashline GLA-RoPE-LH41.819.36.60.50.149.154.846.729.112.50.61.213.727.748.18.721.9
GLA-NoPE-LH57.735.916.57.79.861.552.448.435.220.924.214.525.536.755.020.632.1
+ Log57.735.517.211.710.061.352.648.835.912.627.018.026.436.655.121.032.4
+ wsz=4k57.743.039.432.822.761.652.648.839.835.026.619.839.140.655.232.440.0
+ EME58.043.240.836.131.260.952.349.040.637.835.632.041.944.055.137.243.1
\hdashline GDN-RoPE-LH49.919.96.62.31.553.252.043.830.115.85.94.016.129.349.710.823.8
GDN-NoPE-LH52.039.129.515.914.357.056.153.240.229.421.118.230.139.354.626.035.5
+ Log51.939.533.022.015.957.156.153.241.835.231.224.332.542.754.630.438.4
+ wsz=4k51.941.238.735.029.257.155.953.044.435.829.522.339.242.654.534.541.2
+ EME51.941.840.541.935.957.156.152.946.340.738.533.642.446.554.539.944.8

License

Author: Jaehun Ryu

Link: https://jaehun.me/en/posts/paper-2610-10114v1/

License: CC BY 4.0

This work is licensed under the Creative Commons Attribution 4.0 International License. You are free to use it for any purpose, including commercial use, as long as you provide proper attribution.

Comments