![[Paper Review] Hardware-Efficient Attention for Fast Decoding](https://moonlight-paper-snapshot.s3.ap-northeast-2.amazonaws.com/arxiv/hardware-efficient-attention-for-fast-decoding-1.png)
[Paper Review] Hardware-Efficient Attention for Fast Decoding
Paper GTA & GLA: Hardware-Efficient Attention That Breaks the ‘Memory-Dominated’ DecodeTL;DRGTA (key–value tying) and GLA …
35 min
Attention Optimization
Inference Acceleration
KV-Cache Optimization
Tensor Parallelism
Long-Context Decoding
![[Paper Review] Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding](https://www.storagereview.com/wp-content/uploads/2025/07/image2-2-png-e1752234784623.webp)