SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
Paper SAS: A Simple Attention Sparsification That Learns Context Ranking Directly Without DistillationTL;DR — In post-training attention …
Paper SAS: A Simple Attention Sparsification That Learns Context Ranking Directly Without DistillationTL;DR — In post-training attention …
Paper VestigeKV: The Degenerate RoPE Branch Carries Its Own Cache Eviction SignalTL;DRIn NoPE-MLA models (Kimi Linear 48B), reading only 11% …
Paper ShallowStream: Index Shallowly, Answer Deeply — Unlocking the Computational Bottleneck of Streaming Video UnderstandingTL;DR — In …
Paper Catch the Moment a Model “Revisits Its Thoughts”: How BeaconKV Compresses Reasoning KV Caches 5.8× with Beacon Queries …
Paper Language Models Control Their Own Attention: Declarative AttentionTL;DR — Existing sparse attention methods still paid $O(N)$ per …
Paper Same Request, Different Answer: Prefix Caching Makes Serving Non-Reproducible, and Quantization Amplifies It TL;DR — Prefix caching is …
Paper SGD-KV: Finding ‘heads that are good at summarizing’ cuts the 1M-token KV cache by up to 75% TL;DR — Attention heads do …
Paper The Scores Were a Waste: The Paradox Where ‘Random Sampling’ Ties the Strongest in KV Cache Eviction — An In-Depth Review …