Hardware-Aware FP4 FlashAttention-4
Paper Why Don’t FP4 Tensor Cores Make Attention Faster? Blackwell FlashAttention-4, Solved with Direct-P and Quantized BackpropTL;DR — …
34 min
FlashAttention
Machine Learning
Paper Why Don’t FP4 Tensor Cores Make Attention Faster? Blackwell FlashAttention-4, Solved with Direct-P and Quantized BackpropTL;DR — …
Paper The Recurrent Half Is the Easy-to-Quantize Half: Why Gated DeltaNet Survives at 4 BitsTL;DR — The 48 recurrent Gated DeltaNet (GDN) …