![[Paper Review] Hardware-Efficient Attention for Fast Decoding](https://moonlight-paper-snapshot.s3.ap-northeast-2.amazonaws.com/arxiv/hardware-efficient-attention-for-fast-decoding-1.png)
[Paper Review] Hardware-Efficient Attention for Fast Decoding
Paper GTA & GLA: Hardware-Efficient Attention That Breaks the ‘Memory-Dominated’ DecodeTL;DRGTA (key–value tying) and GLA …
All posts on technology, daily life, and thoughts.
![[Paper Review] Hardware-Efficient Attention for Fast Decoding](https://moonlight-paper-snapshot.s3.ap-northeast-2.amazonaws.com/arxiv/hardware-efficient-attention-for-fast-decoding-1.png)
Paper GTA & GLA: Hardware-Efficient Attention That Breaks the ‘Memory-Dominated’ DecodeTL;DRGTA (key–value tying) and GLA …
![[Paper Review] Pretraining Large Language Models with NVFP4](https://developer-blogs.nvidia.com/wp-content/uploads/2025/08/Optimizing-LLM-Training-png.webp)
Paper NVFP4 4-bit Pretraining, Made Practical: 12B up to 10T Tokens, Effectively On Par with FP8TL;DRPretraining a 12B hybrid …
Paper “Let’s put an OS on the GPU”: A proposal for a GPU multitasking OS layer for the LLM eraOne-line summary (TL;DR)In …
Paper Link DroidSpeak: Reducing Prefill Latency by 1.7–3.1× through Cross-LLM Prefix-KV ReuseTL;DRWhen multiple LLMs share the same …
![[Paper Review] Marconi: Prefix Caching for the Era of Hybrid LLMs](https://pbs.twimg.com/media/GdyLXO9W4AADox0.jpg)
Paper Link Marconi: Rethinking Prefix Caching for the Hybrid LLM EraTL;DRMarconi introduces a prefix-caching framework for hybrid LLM …
![[Paper Review] SGLang: Efficient Execution of Structured Language Model Programs](https://cdn.bytez.com/mobilePapers/v2/neurips/94872/images/20-0.png)
Paper Link SGLang & RadixAttention: How Execution Optimization for “LM Programs” Achieved a 6.4x SpeedupTL;DRBy combining a …