SlimKV: Compressing Tokens and Features Together, Cutting the KV Cache by up to 32× with Reconstruction-Free Latent-Space Attention
TL;DR
SlimKV is a question-agnostic framework that compresses the KV cache — the bottleneck of long-context LLM inference — simultaneously along the token axis and the feature axis. The key idea is to remove RoPE only from the Keys of beacon memory tokens, so attention is performed directly in the latent space without a reconstruction step. It achieves the best performance at 16×/32× compression on LongBench (source: §5.2), and at a 128K token length it delivers up to 7.34× attention speedup and 3.38× end-to-end decoding speedup over no compression (source: §5.4).
Core Idea
The central claim of this paper can be summed up in one sentence.
The authors hypothesize that by training the beacon state in low-rank form from the start and removing RoPE only from the beacon Key, one can achieve 16–32× KV cache compression that simultaneously overcomes the limits of single-axis token/feature compression and the Key reconstruction latency bottleneck.
Three lines of contribution support this claim (source: §1).
- Combined token·feature compression framework — It trains low-rank beacon KV representations within a fixed feature budget and allocates adaptive ranks to match per-layer heterogeneity.
- Discovery of positional asymmetry — It empirically shows that removing RoPE on the Key side (K-RoPE) damages ordinary tokens and beacon tokens differently, with beacons being far less sensitive.
- Reconstruction-free latent-space attention (LSA) — Using this asymmetry, it removes RoPE only from the beacon Key, thereby skipping the restoration to full dimensionality during decoding.
Key Numbers at a Glance
| Item | Value |
|---|---|
| Backbone models | Llama-3.1-8B-Instruct / Qwen2.5-14B-Instruct / Qwen3-30B-A3B-Instruct (MoE) |
| Compression scheme | token × feature combined (e.g. t4f4 = 4× token · 4× feature = 16×) |
| LongBench uncompressed average | 49.50 (Llama-8B) / 51.52 (Qwen-14B) / 50.30 (Qwen3-MoE) |
| 16× compression average (SlimKV-f4) | 46.36 (Llama-8B) / 46.23 (Qwen-14B) |
| Low-compression (4×/8×) retention | over 96% of uncompressed |
| Max attention speedup (128K, 16×) | 7.34× |
| Max E2E speedup (128K, 16×) | 3.38× |
| Trainable parameters | 621M (Llama-8B) / 1.41B (Qwen-14B) / 434M (Qwen3-MoE, 1.45% of total) |
Background: The Problem They Solved
LLMs need long context for document understanding, multi-turn dialogue, and complex reasoning, but during inference the KV cache grows in proportion to context length, causing memory pressure and memory-access latency (source: §1). The KV cache size is roughly
$$ \text{KV-Cache(GB)} \approx \frac{2 \cdot L \cdot H \cdot d_{\text{head}} \cdot \text{seq} \cdot \text{batch} \cdot \text{bytes/elt}}{10^9} $$where $L$ is the number of layers, $H$ is the number of KV heads, $d_{\text{head}}$ is the per-head dimension, seq is the sequence length, and batch is the batch size. The longer the context, the more this term dominates memory and bandwidth.
Existing compression strategies fall broadly into two axes, each with clear limits (source: §2).
- Token-wise: reduces the number of cache states. But explicit eviction loses a lot of information under aggressive compression (KVZip), and most eviction methods must first build the entire KV before eviction, which limits prefill compute savings (SnapKV, H2O). Soft compression (Activation Beacon) mitigates information loss but did not further explore the redundancy of the beacon state.
- Feature-wise: reduces the KV dimension per token (MQA, GQA, MLA, Eigen Attention, PALU, MHA2MLA). However, the compressed Key often has to be reconstructed to full dimensionality before RoPE is applied, leaving a decoding latency bottleneck.
The central question the authors pose is this (source: §1).
“Can combined token·feature compression go beyond the limits of a single axis while also avoiding the reconstruction bottleneck?”
The New Approach: SlimKV
SlimKV is a framework that combines feature-axis compression on top of an Activation Beacon-style progressive token-compression pipeline (source: §3.1). The workflow is as follows.
- Split the input into chunks at fixed intervals, and insert a beacon token between each chunk.
- Let each chunk attend to the previous beacon memories so the beacon state absorbs context information.
- When a chunk’s prefill finishes, discard the KV of the raw tokens and cache only the compressed beacon memory.
SlimKV differs in two ways.
(1) Low-Rank Beacon Training — Instead of post-hoc SVD decomposition, it trains the beacon KV projection in the target low-rank form from the start (source: §3.2). The backbone is frozen and only the beacon branch is updated. The beacon hidden state $h_b^{(l)}$ (at layer $l$) is projected as
$$ C_K^{(l)} = h_b^{(l)} W_{K,\downarrow}^{(l)},\qquad C_V^{(l)} = h_b^{(l)} W_{V,\downarrow}^{(l)} $$into latent states $C_K^{(l)}, C_V^{(l)}$ smaller than the original KV dimension $D$, and the full-dimensional form is reconstructed as $K_b^{(l)} = C_K^{(l)} U_K^{(l)\top}$, $V_b^{(l)} = C_V^{(l)} U_V^{(l)\top}$. At inference time, only the latent states $C_K, C_V$ are stored in the cache.
(2) Layer-Adaptive Rank Allocation — The low-rank structure of transformer projection matrices varies by layer. Since the beacon projection is unknown before training, the authors use the pretrained original KV projection weights as a proxy (source: §3.3). The two exhibit per-layer rank patterns with a Spearman correlation of about 0.99 or higher and a per-layer rank error within 2, matching almost exactly (source: Appx.A, Tab.SVD).
Specifically, for the singular values $\sigma_{l,1} \ge \cdots \ge \sigma_{l,D}$ of the original KV projection, they define the cumulative spectral energy
$$ E_l(r) = \frac{\sum_{j=1}^{r} \sigma_{l,j}^{2}}{\sum_{j=1}^{D} \sigma_{l,j}^{2}} $$and pick the initial rank $r_l^{0} = \min\{ r \mid E_l(r) \ge \rho \}$ with an energy threshold $\rho$. They then introduce a “floor” $r_{\min} = \lceil \lambda \bar r \rceil$ to keep layers with too-small ranks from becoming a training bottleneck, and choose the largest $\rho$ that satisfies the total budget $B = LD/k$. This paper uses $\lambda = 0.5$ (source: §5.1).
How It Works: A Concrete Example
Let us walk through the whole flow, from compression to speedup, with a small example. To make this easier, assume the following setting (source: §4.4, Appx.C).
- KV heads $H_{kv}=8$, query heads $H_q=32$, GQA group $g = H_q/H_{kv} = 4$
- Per-head dimension $d=128$, total KV dimension $D = H_{kv} \cdot d = 1024$
- Feature compression ratio $k=4$ → latent KV dimension $D/k = 256$
Step 1 — Token compression. Summarizing a 16K-token document into one beacon every 4 tokens yields 4K beacons (4× token compression).
Step 2 — Feature compression. Each beacon’s K/V is projected from $D=1024$ dimensions into a $256$-dimensional latent space (4× feature compression). Multiplying the two gives 16× compression, and the cache retains only $C_K, C_V \in \mathbb{R}^{4000 \times 256}$.
Step 3 — Reconstruction-free attention. Now suppose we compute attention for a single new token during decoding. If RoPE remains on the beacon Key, RoPE is position-dependent and cannot be absorbed into the latent space in advance, so we must first reconstruct the full dimensions via $K_b = C_K U_K^\top$ and then apply RoPE.
$$ Q' \operatorname{RoPE}(K_b)^\top = Q' \operatorname{RoPE}(C_K U_K^\top)^\top $$In contrast, SlimKV has removed RoPE from the beacon Key, so using the associativity of matrix multiplication it computes the score directly in the latent space (source: §4.3).
$$ Q' K_b^\top = Q'(C_K U_K^\top)^\top = (Q' U_K) C_K^\top $$The Value path is handled the same way, $A V_b = A (C_V U_V^\top) = (A C_V) U_V^\top$. In other words, the two steps of “reconstruction → attention” become the single step of “query projection → latent attention.”
How much does the cost change? Comparing the dominant FLOPs of one decoding step (source: §4.4, Appx.C)
$$ F_{\text{recon}} \approx N_b \frac{2D^2}{k} + (1+2g) N_b D $$$$ F_{\text{SlimKV}} \approx \frac{2gD^2}{k} + \frac{2H_q}{k} N_b D $$where $N_b$ is the number of cached beacon tokens. The reconstruction approach incurs a full-dimensional reconstruction cost of $\mathcal{O}(N_b D^2/k)$ proportional to $N_b$, whereas SlimKV turns that term into $\mathcal{O}(gD^2/k)$, which applies only to the current step’s query and output projections, independent of $N_b$. Plugging in the setting above,
$$ F_{\text{recon}} \approx \tfrac{1}{2}N_b D^2 + 9 N_b D \approx 521 N_b D,\qquad F_{\text{SlimKV}} \approx 2D^2 + 16 N_b D \approx 16 N_b D $$so for long contexts ($N_b \gg 128$) the theoretical arithmetic is reduced by about 32.6×. Of course, this accounts only for arithmetic cost and excludes kernel overhead and memory bandwidth, so the measured speedup is more conservative, as shown below (source: Appx.C).
Performance Validation: Main Results
LongBench — The More Compression, the Better
In the high-compression 16×/32× regime, SlimKV set the best average score on both backbones (source: §5.2, Tab.1). Notably, SlimKV-f4 (t4f4, 4× token · 4× feature) at 16× compression reaches Llama-3.1-8B 46.36 and Qwen2.5-14B 46.23, beating SnapKV, Activation Beacon, PALU, and KVZip. At 32× compression (t8f4) it is also best at 44.75 / 45.09 respectively.
Comparing feature compression ratios of 2×/4×/8×, 4× was the sweet spot giving the best average in most high-compression settings. This is due to the tradeoff: too little feature compression makes token compression excessive, while too much leaves insufficient latent capacity (source: §5.2).
Comparing the gap from uncompressed against Activation Beacon, SlimKV-f4 narrows the gap by 18.5%–28.2% (source: Tab.2). At low compression (4×/8×) it uses 2× feature compression (SlimKV-f2) to retain over 96% of the uncompressed score, and on Qwen2.5-14B at 8× reaches 49.94 (uncompressed 51.52), far ahead of KVZip and PALU (source: Tab.3).
Question-Independent vs Question-Aware: The Strength Revealed in Multi-Turn
Question-aware methods like SnapKV benefit from knowing the question in advance. SlimKV is question-independent, so it shines in scenarios that reuse a single compressed context across multiple questions (source: §5.2). In the QMSum multi-request experiment, against an uncompressed ROUGE-L F1 of 25.17, SnapKV plummets to 10.39 at 16× compression, while SlimKV-f4 holds 24.47, 97.2% of uncompressed.
Needle-in-a-Haystack — Positional Robustness Survives Even Without K-RoPE
Removing RoPE from the beacon Key raises concerns about losing positional information. The NIAH results put this concern to rest (source: §5.2, Fig.7). SlimKV-f4 showed stable retrieval performance across needle depth, and maintained performance even when the answer lay beyond Qwen2.5-14B’s context window (32K). This is because long context is compressed into a small number of beacon tokens, keeping the cache within the usable range, and the positional cues absorbed during RoPE-aware prefill remain valid even without K-RoPE.
Efficiency — The Real Benefit of Eliminating Reconstruction
Single-step attention and end-to-end latency were measured on a single A100 80GB, BF16, batch size 1, with Llama-3.1-8B (source: §5.4, Tab.5). Compared with the reconstruction-based variant (Recon, which keeps beacon K-RoPE and reconstructs full dimensions), SlimKV’s speedup grows as context lengthens.
| Compression | Metric | 32K | 64K | 96K | 128K |
|---|---|---|---|---|---|
| 16× | Attention speedup | 1.94× | 3.81× | 5.08× | 7.34× |
| 16× | E2E speedup | 1.22× | 2.04× | 2.63× | 3.38× |
| 32× | Attention speedup | 1.62× | 2.54× | 4.36× | 6.09× |
| 32× | E2E speedup | 1.23× | 1.75× | 2.82× | 3.49× |
Also, at the maximum length of 11.5K where the uncompressed model can run without OOM, SlimKV improves TTFT by 1.03–1.37× and reduces peak memory by 2.71–2.82× (source: §5.4). Interestingly, the peak memory of Recon and SlimKV is nearly the same. Removing reconstruction changes the compute flow, not the storage size (source: Appx.F).
Scalability — It Transfers Well Even to a 30B MoE
On Qwen3-30B-A3B-Instruct (MoE), SlimKV-f4 maintained competitive performance at 46.08 for 16× and 45.60 for 32× (uncompressed 50.30) (source: §5.3, Tab.4). Its trainable parameters are only 434M, just 1.45% of the full model, demonstrating the potential for parameter-efficient scaling.
Ablation — The Fact That “No Reconstruction” Is Not Free
The 32× compression ablation is especially telling (source: §5.5, Tab.6). Simply removing beacon K-RoPE drops performance from 43.67 to 41.28. Removing RoPE alone is not enough — one must train directly in low-rank form (41.95 → 44.38) and add layer-adaptive rank allocation (44.75) to surpass Activation Beacon.
Our Perspective: Strengths, Limits, and Why This Research Matters
Strengths. The most valuable contribution of this work is that it redraws the problem not as a single axis of “accuracy vs efficiency” but as a structure of two axes of compression × the reconstruction bottleneck. In particular, (1) the empirical discovery of beacon K-RoPE asymmetry and (2) the logical development that connects it to eliminating reconstruction are clean. Even on the attention heatmap similarity metric, removing beacon K-RoPE gives a Pearson correlation of 0.9707 with the original pattern, far closer than 0.7570 for removing original tokens (source: §4.2, Tab.5). The fact that the cost analysis honestly separates theory (32.6×) from measurement (up to 7.34×) also inspires confidence.
Limits. The authors themselves acknowledge two (source: §Limitations). First, validation is confined to RoPE-based LLMs. Second, it is text-only, and extension to multimodal/cross-modal KV caches is unverified. From a reader’s perspective, we would add that (a) requiring training is a barrier to adoption compared with training-free methods like PALU and KVZip (though the number of trainable parameters is small), (b) having erased the RoPE of the beacon Key, there is considerable backbone-to-backbone variation on order-sensitive tasks (especially code and format-dependent tasks) (source: Appx.G), and (c) the NIAH judgment relies on automatic GPT-5.4-mini grading, which caps the robustness of the measurement.
Why this research matters. In long-context serving, the KV cache is already the main memory and bandwidth bottleneck, and compression techniques so far have failed to simultaneously address “how well is accuracy preserved” and “does it actually get faster.” SlimKV proposes a design principle that answers both questions at once — beacons do not need to express position via RoPE — and this is likely to become a baseline reference for future low-rank KV design.
What’s Next?: The Road Ahead
The authors themselves suggest two directions (source: §Limitations). One is verifying generalization to positional encodings other than RoPE (e.g., ALiBi, YaRN, learned positions). The paper theoretically shows that ALiBi is algebraically compatible as $S = Q K_b^\top + B = (Q U_K) C_K^\top + B$ (source: Appx.J), so empirical validation is the natural next step. The other is extension to visual/cross-modal KV caches.
Adding further, the reasonably expected follow-up work can be summarized as follows.
- Shoring up order-sensitive tasks: Injecting a subtle positional cue (e.g., a lightweight ALiBi-style bias) to offset the positional information loss of the beacon Key could reduce backbone-to-backbone variation on code and format tasks.
- Quantitative comparison with MLA: Although the appendix gives a conceptual comparison (source: Appx.I), a head-to-head comparison against MLA-style low-rank designs at the same backbone and same compression ratio is still lacking.
- System-level optimization: A kernel implementation for latent-space attention is the key lever that could narrow the gap between the measured speedup (7.34×) and the theoretical speedup (32.6×).
- Reducing training cost: Currently training takes hours to tens of hours per backbone, so research on initialization and regularization that converges with less corpus would lower the deployment barrier.
SlimKV is a study that combines practicality and insight, turning the conventional wisdom that “compression means loss” into “where you compress determines the loss.”
Tables from the paper
Tables converted mechanically from the arXiv e-print LaTeX source. The numbers are the paper’s own and did not pass through a model.
Table 1. Layer-wise rank consistency between Beacon and raw K/V projections.
| Model | Proj. | CR | Beacon | Raw | Spearman | MAE / Max |
|---|---|---|---|---|---|---|
| Llama | K | 2 | 64.00 | 64.00 | 1.000 | 0.00 / 0 |
| K | 4 | 32.00 | 32.00 | 0.996 | 0.06 / 1 | |
| K | 8 | 16.00 | 16.00 | 0.996 | 0.06 / 1 | |
| V | 2 | 64.00 | 64.00 | 0.990 | 0.50 / 1 | |
| V | 4 | 32.00 | 32.00 | 0.954 | 0.50 / 2 | |
| V | 8 | 16.00 | 16.00 | 0.928 | 0.38 / 2 | |
| Qwen | K | 2 | 64.00 | 64.00 | 0.997 | 0.08 / 1 |
| K | 4 | 32.00 | 32.00 | 0.999 | 0.04 / 1 | |
| K | 8 | 16.00 | 16.00 | 0.997 | 0.04 / 1 | |
| V | 2 | 64.00 | 64.00 | 0.997 | 0.08 / 1 | |
| V | 4 | 32.00 | 32.00 | 1.000 | 0.00 / 0 | |
| V | 8 | 16.00 | 16.00 | 0.990 | 0.04 / 1 |
Table 2. Detailed decoding latency on Llama-3.1-8B-Instruct over 128 generated tokens. Attn. and E2E denote cumulative attention and end-to-end decoding latency. All values are in seconds.
| Comp. | Method | 32K Attn. | 32K E2E | 64K Attn. | 64K E2E | 96K Attn. | 96K E2E | 128K Attn. | 128K E2E |
|---|---|---|---|---|---|---|---|---|---|
| 1$\times$ | Full KV | 11.84 | 13.28 | 23.97 | 25.40 | 34.93 | 36.32 | 46.64 | 48.08 |
| 16$\times$ | Recon | 6.74 | 11.75 | 8.52 | 15.46 | 9.63 | 17.02 | 8.73 | 16.79 |
| SlimKV | 6.11 | 10.93 | 6.29 | 12.47 | 6.87 | 13.79 | 6.35 | 14.24 | |
| 32$\times$ | Recon | 7.80 | 11.56 | 10.18 | 15.23 | 9.11 | 13.95 | 9.48 | 15.69 |
| SlimKV | 7.31 | 10.81 | 9.43 | 14.51 | 8.01 | 12.90 | 7.65 | 13.78 |
Table 3. TTFT and peak allocated memory on Llama-3.1-8B-Instruct. TTFT is reported in seconds and memory in GB.
| Comp. | Method | 32K | 64K | 96K | 128K |
|---|---|---|---|---|---|
| TTFT (s) | |||||
| 16$\times$ | Recon | 16.59 | 46.38 | 89.15 | 153.74 |
| SlimKV | 16.16 | 44.80 | 87.12 | 152.57 | |
| 32$\times$ | Recon | 11.28 | 28.34 | 50.13 | 82.73 |
| SlimKV | 10.92 | 27.77 | 48.62 | 79.96 | |
| Peak Memory (GB) | |||||
| 16$\times$ | Recon | 25.73 | 29.77 | 33.80 | 37.84 |
| SlimKV | 25.72 | 29.76 | 33.80 | 37.84 | |
| 32$\times$ | Recon | 23.42 | 25.26 | 27.09 | 28.92 |
| SlimKV | 23.41 | 25.25 | 27.08 | 28.91 |
Table 4. Similarity to the w/ K-RoPE setting.
| Variant | Attn. $\uparrow$ | Block $\uparrow$ | Lag JS $\downarrow$ | Col. JS $\downarrow$ |
|---|---|---|---|---|
| Beacon removed | 0.9707 | 0.9982 | 0.0002 | 0.0030 |
| Raw removed | 0.7570 | 0.8922 | 0.0062 | 0.0349 |
Table 5. LongBench scores under high KV-cache compression. SlimKV t$n$f$m$: $n\times$ token-wise, $m\times$ feature-wise. The bold and underlined numbers denote the first and second rankings, respectively.
| Comp. | Method | Llama-3.1-8B S-Doc | Llama-3.1-8B M-Doc | Llama-3.1-8B Summ. | Llama-3.1-8B Few | Llama-3.1-8B Code | Llama-3.1-8B Avg. | Qwen2.5-14B S-Doc | Qwen2.5-14B M-Doc | Qwen2.5-14B Summ. | Qwen2.5-14B Few | Qwen2.5-14B Code | Qwen2.5-14B Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1$\times$ | Full KV | 44.34 | 46.65 | 28.71 | 69.45 | 58.37 | 49.50 | 42.72 | 43.73 | 28.05 | 72.04 | 61.05 | 51.52 |
| 16$\times$ | KVZip | 33.86 | 29.11 | 23.97 | 22.97 | 34.89 | 28.96 | 7.83 | 5.48 | 3.27 | 35.12 | 14.20 | 13.18 |
| PALU | 13.18 | 9.80 | 12.75 | 37.86 | 21.04 | 18.92 | 17.77 | 14.18 | 21.40 | 59.40 | 25.75 | 27.70 | |
| SnapKV | 38.55 | 43.30 | 23.61 | 65.92 | 56.62 | 45.60 | 36.74 | 51.20 | 20.22 | 68.50 | 52.37 | 45.81 | |
| Activation Beacon | 34.72 | 36.28 | 25.34 | 65.01 | 65.14 | 45.30 | 35.76 | 38.31 | 25.24 | 63.66 | 57.77 | 44.15 | |
SlimKV-f2 (t8f2) | 37.64 | 37.40 | 26.25 | 65.11 | 64.55 | 46.19 | 39.44 | 42.16 | 25.91 | 67.20 | 57.72 | 46.48 | |
SlimKV-f4 (t4f4) | 38.65 | 34.33 | 26.57 | 65.14 | 67.11 | 46.36 | 38.15 | 43.76 | 25.50 | 69.16 | 54.61 | 46.23 | |
SlimKV-f8 (t2f8) | 33.46 | 36.24 | 26.02 | 65.68 | 63.82 | 45.05 | 38.25 | 38.83 | 25.61 | 68.06 | 50.89 | 44.33 | |
| 32$\times$ | KVZip | 9.50 | 3.35 | 12.69 | 0.96 | 20.48 | 9.40 | 1.09 | 0.78 | 1.32 | 1.53 | 20.91 | 5.12 |
| PALU | 1.58 | 0.56 | 5.47 | 1.32 | 15.02 | 4.79 | 4.20 | 2.82 | 10.90 | 14.66 | 13.67 | 9.25 | |
| SnapKV | 35.80 | 42.82 | 21.61 | 62.70 | 53.29 | 43.24 | 32.45 | 49.61 | 18.09 | 62.71 | 48.54 | 42.28 | |
| Activation Beacon | 29.94 | 35.24 | 24.53 | 64.26 | 64.35 | 43.67 | 32.80 | 35.64 | 25.16 | 63.20 | 57.01 | 42.76 | |
SlimKV-f2 (t16f2) | 33.49 | 35.34 | 24.86 | 61.19 | 59.68 | 42.91 | 37.15 | 38.59 | 25.49 | 64.97 | 57.09 | 44.66 | |
SlimKV-f4 (t8f4) | 34.16 | 35.75 | 25.43 | 63.35 | 65.07 | 44.75 | 37.55 | 40.29 | 25.50 | 68.76 | 53.36 | 45.09 | |
SlimKV-f8 (t4f8) | 32.53 | 35.41 | 25.14 | 64.36 | 62.22 | 43.93 | 35.83 | 37.99 | 25.27 | 68.27 | 56.10 | 44.69 |
Table 6. LongBench gap reduction over Activation Beacon. $\Delta$ denotes the drop from Full KV.
| Backbone | Comp. | $\Delta_{\textbf{AB}}$ | $\Delta_{\textbf{SlimKV-f4}}$ | $\Delta$ Red. | Gap Red. |
|---|---|---|---|---|---|
| Llama-3.1-8B | 16$\times$ | 4.20 | 3.14 | 1.06 | 25.2% |
| 32$\times$ | 5.83 | 4.75 | 1.08 | 18.5% | |
| Qwen2.5-14B | 16$\times$ | 7.37 | 5.29 | 2.08 | 28.2% |
| 32$\times$ | 8.76 | 6.43 | 2.33 | 26.6% |
Table 7. LongBench averages at low compression. The averages of the uncompressed models are in Table .
| Backbone | Comp. | KVZip | PALU | AB | SlimKV-f2 |
|---|---|---|---|---|---|
| Llama-3.1-8B | 4$\times$ | 36.39 | 46.66 | 48.21 | 48.10 |
| 8$\times$ | 37.25 | 46.42 | 46.93 | 47.77 | |
| Qwen2.5-14B | 4$\times$ | 47.72 | 45.95 | 48.86 | 49.92 |
| 8$\times$ | 46.74 | 45.15 | 45.50 | 49.94 |
Table 8. LongBench category scores on Qwen3-30B-A3B-Instruct-2507.
| Method | Comp. | S-Doc | M-Doc | Summ. | Few-shot | Code | Avg. |
|---|---|---|---|---|---|---|---|
| Full KV | 1$\times$ | 42.87 | 51.02 | 25.60 | 71.26 | 60.75 | 50.30 |
| SlimKV-f4 | 16$\times$ | 36.87 | 40.43 | 26.00 | 68.24 | 58.87 | 46.08 |
| SlimKV-f4 | 32$\times$ | 34.14 | 38.16 | 24.77 | 67.85 | 63.09 | 45.60 |
Comments