Hardware-Aware FP4 FlashAttention-4
Paper Why Don’t FP4 Tensor Cores Make Attention Faster? Blackwell FlashAttention-4, Solved with Direct-P and Quantized BackpropTL;DR — …
34 min
FlashAttention
Machine Learning
Paper Why Don’t FP4 Tensor Cores Make Attention Faster? Blackwell FlashAttention-4, Solved with Direct-P and Quantized BackpropTL;DR — …
![[Paper Review] Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality](https://icml.cc/media/PosterPDFs/ICML%202024/32613.png)
Paper Link Structured State Space Duality: Unifying SSMs and Attention with Mamba-2 for 2–8× AccelerationTL;DRStructured State-Space Duality …
Paper Link Hydragen: The Secret Weapon for Decoding Large Batches with Shared Prefixes up to 32× FasterTL;DRBy decomposing the prefix and …
![[Paper Review] Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding](https://www.storagereview.com/wp-content/uploads/2025/07/image2-2-png-e1752234784623.webp)
Paper Link Helix Parallelism: Breaking the Latency-Throughput Wall of Ultra-Long LLM DecodingTL;DRHelix Parallelism schedules Attention and …
Paper Native Sparse Attention (NSA) — 11× faster even at 64k tokens, accuracy intactOne-line summary (TL;DR)NSA combines a three-branch …