Language Models Can Control Their Own Attention
Paper Language Models Control Their Own Attention: Declarative AttentionTL;DR — Existing sparse attention methods still paid $O(N)$ per …
14 min
KV Cache
Chain-of-Thought
Long Context
Natural Language Processing