![[Paper Review] Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models](https://paper-assets.alphaxiv.org/figures/2601.07372v1/img-0.jpeg)
[Paper Review] Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
Paper Engram: The Second Sparsity Axis After Conditional Computation (MoE), Conditional MemoryOne-Line Summary (TL;DR)Engram adds a …
![[Paper Review] Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models](https://paper-assets.alphaxiv.org/figures/2601.07372v1/img-0.jpeg)
Paper Engram: The Second Sparsity Axis After Conditional Computation (MoE), Conditional MemoryOne-Line Summary (TL;DR)Engram adds a …
![[Paper Review] Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin](https://cdn-uploads.huggingface.co/production/uploads/6317233cc92fd6fee317e030/yNP71PjobVvLDgJ0R0qV2.png)
Paper One Narrative Forged by Massive Activations: Bridging Attention Sink and Compression ValleyTL;DRMassive activations in the residual …
![[Paper Review] Hardware-Efficient Attention for Fast Decoding](https://moonlight-paper-snapshot.s3.ap-northeast-2.amazonaws.com/arxiv/hardware-efficient-attention-for-fast-decoding-1.png)
Paper GTA & GLA: Hardware-Efficient Attention That Breaks the ‘Memory-Dominated’ DecodeTL;DRGTA (key–value tying) and GLA …
Paper “Let’s put an OS on the GPU”: A proposal for a GPU multitasking OS layer for the LLM eraOne-line summary (TL;DR)In …
![[Paper Review] Marconi: Prefix Caching for the Era of Hybrid LLMs](https://pbs.twimg.com/media/GdyLXO9W4AADox0.jpg)
Paper Link Marconi: Rethinking Prefix Caching for the Hybrid LLM EraTL;DRMarconi introduces a prefix-caching framework for hybrid LLM …
![[Paper Review] SGLang: Efficient Execution of Structured Language Model Programs](https://cdn.bytez.com/mobilePapers/v2/neurips/94872/images/20-0.png)
Paper Link SGLang & RadixAttention: How Execution Optimization for “LM Programs” Achieved a 6.4x SpeedupTL;DRBy combining a …
![[Paper Review] Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality](https://icml.cc/media/PosterPDFs/ICML%202024/32613.png)
Paper Link Structured State Space Duality: Unifying SSMs and Attention with Mamba-2 for 2–8× AccelerationTL;DRStructured State-Space Duality …
Link to Paper Dynamic Memory Sparsification (DMS): Making LLM Hyper-Scaling a Reality with 8× KV Cache CompressionOne-Line Summary …
Paper Link Hydragen: The Secret Weapon for Decoding Large Batches with Shared Prefixes up to 32× FasterTL;DRBy decomposing the prefix and …
![[Paper Review] KIMI K2: OPEN AGENTIC INTELLIGENCE](https://github.com/MoonshotAI/Kimi-K2/raw/main/figures/kimi-logo.png)
Paper Link Kimi K2: An Open-Source LLM’s Leap Toward Agentic IntelligenceTL;DRWith a 3-stage pipeline consisting of MuonClip pretraining + …
Paper Link Qwen 3: The Evolution of a Giant MoE Language Model with Adjustable Reasoning DepthTL;DR (in one line)Qwen 3 couples a …
![[Paper Review] Massive Activations in Large Language Models](https://eric-mingjie.github.io/massive-activations/assets/main_teaser_final.png)
Paper Link Massive Activations, Hidden Biases: A Reinterpretation of Self-Attention’s Secrets TL;DRJust 4–10 extreme scalar values (×10,000) …
Paper Link Peri-LayerNorm: A Third Option Beyond Post-LN and Pre-LNTL;DRBy simply adding another LayerNorm right after the residual …
![[paper review] SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-bit Training](https://cdn-uploads.huggingface.co/production/uploads/66c0a08bac74db25de8427ec/Tb20E3IJSV6PjcD9Nkvfg.png)
[Paper Review] SageAttention 3 & SageBwd — FP4-Powered Inference and 8-bit TrainingPaper link: https://arxiv.org/abs/2505.11594v1 📝 …
![[Paper Review] Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding](https://www.storagereview.com/wp-content/uploads/2025/07/image2-2-png-e1752234784623.webp)
Paper Link Helix Parallelism: Breaking the Latency-Throughput Wall of Ultra-Long LLM DecodingTL;DRHelix Parallelism schedules Attention and …
Paper Subgoal Curriculum + CoT Consistency: DeepSeek-Prover-V2 Reshapes Automated Theorem ProvingTL;DRDeepSeek-Prover-V2, built on the …
Paper Inference-Time Scaling: How DeepSeek-GRM Surpassed Giant ModelsOne-Line Summary (TL;DR)“27B model × 32 samples”—With only …
Paper CODE I/O: From Code I/O + Natural-Language CoT to General-Purpose Reasoning — Lifting 7B-30B LLMs by +2 Points on Average with Data …
Paper Native Sparse Attention (NSA) — 11× faster even at 64k tokens, accuracy intactOne-line summary (TL;DR)NSA combines a three-branch …
Paper Janus-Pro 7B: Dual-Encoder Multimodal LLM That Outsmarts Bigger ModelsOne-line summary (TL;DR)By fully separating the SigLIP …
Paper One-line Summary (TL;DR)DeepSeek-V3 is an open-source SOTA that combines Aux-loss-free Load-Balancing Bias, FP8 mixed-precision …
Paper One-line summary (TL;DR)JanusFlow achieves state-of-the-art performance in both image understanding and generation (FID 9.51 / GenEval …