ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding
Paper ShallowStream: Index Shallowly, Answer Deeply — Unlocking the Computational Bottleneck of Streaming Video UnderstandingTL;DR — In …
Paper ShallowStream: Index Shallowly, Answer Deeply — Unlocking the Computational Bottleneck of Streaming Video UnderstandingTL;DR — In …
Paper Uno: AR and Diffusion in One Model — Lossless Parallel Generation for LLM AccelerationTL;DR — LLMs are slow because next-token …
Paper Catch the Moment a Model “Revisits Its Thoughts”: How BeaconKV Compresses Reasoning KV Caches 5.8× with Beacon Queries …
Paper Why Don’t FP4 Tensor Cores Make Attention Faster? Blackwell FlashAttention-4, Solved with Direct-P and Quantized BackpropTL;DR — …
Paper Language Models Control Their Own Attention: Declarative AttentionTL;DR — Existing sparse attention methods still paid $O(N)$ per …
Paper Same Request, Different Answer: Prefix Caching Makes Serving Non-Reproducible, and Quantization Amplifies It TL;DR — Prefix caching is …
Paper SGD-KV: Finding ‘heads that are good at summarizing’ cuts the 1M-token KV cache by up to 75% TL;DR — Attention heads do …
Paper SMELT: Scaling Laws for Compute-Matched MoE Loop TransformersTL;DRLoop transformers scale depth by executing a layer block repeatedly, …
Paper The Recurrent Half Is the Easy-to-Quantize Half: Why Gated DeltaNet Survives at 4 BitsTL;DR — The 48 recurrent Gated DeltaNet (GDN) …