![[Paper Review] Hardware-Efficient Attention for Fast Decoding](https://moonlight-paper-snapshot.s3.ap-northeast-2.amazonaws.com/arxiv/hardware-efficient-attention-for-fast-decoding-1.png)
[Paper Review] Hardware-Efficient Attention for Fast Decoding
Paper GTA & GLA: Hardware-Efficient Attention That Breaks the ‘Memory-Dominated’ DecodeTL;DRGTA (key–value tying) and GLA …
35 min
Attention Optimization
Inference Acceleration
KV-Cache Optimization
Tensor Parallelism
Long-Context Decoding
![[paper review] SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-bit Training](https://cdn-uploads.huggingface.co/production/uploads/66c0a08bac74db25de8427ec/Tb20E3IJSV6PjcD9Nkvfg.png)