Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference
Paper Cache-Learning Routers: Cache-Aware Joint Router Adaptation for Memory-Constrained MoE InferenceTL;DRStarting from the observation …
All posts on technology, daily life, and thoughts.
Paper Cache-Learning Routers: Cache-Aware Joint Router Adaptation for Memory-Constrained MoE InferenceTL;DRStarting from the observation …
Paper VestigeKV: The Degenerate RoPE Branch Carries Its Own Cache Eviction SignalTL;DRIn NoPE-MLA models (Kimi Linear 48B), reading only 11% …
Paper ShallowStream: Index Shallowly, Answer Deeply — Unlocking the Computational Bottleneck of Streaming Video UnderstandingTL;DR — In …
Paper Uno: AR and Diffusion in One Model — Lossless Parallel Generation for LLM AccelerationTL;DR — LLMs are slow because next-token …
Paper Catch the Moment a Model “Revisits Its Thoughts”: How BeaconKV Compresses Reasoning KV Caches 5.8× with Beacon Queries …
Paper Why Don’t FP4 Tensor Cores Make Attention Faster? Blackwell FlashAttention-4, Solved with Direct-P and Quantized BackpropTL;DR — …