Kalman Delta Networks: Uncertainty-aware Associative Memory
Linear Attention Rewritten with a Kalman Filter: Kalman Delta NetworksTL;DRThe fixed-size recurrent memory update of linear attention is …
All posts on technology, daily life, and thoughts.
Linear Attention Rewritten with a Kalman Filter: Kalman Delta NetworksTL;DRThe fixed-size recurrent memory update of linear attention is …
Paper X-AuT: Pruning Audio Encoders in Speech LLMs with Behavioral Probes and Cross-Scale DistillationTL;DRX-AuT is a framework that …
Paper Deadline-Filling Prefill Chunks: SLOWeave — An Adaptive Chunking Scheduler for LLM Serving TL;DR — Instead of a fixed prefill chunk …
Paper Why Are Video LLMs Still So Expensive? — A Four-Stage Guide to Inference Efficiency MechanismsTL;DR — This paper classifies 125 …
Paper Miles v0.1: A “verified, clean, and scalable” full-stack system for frontier RL post-trainingTL;DR: Miles is a full-stack …
Paper Two Designs That Make Speculative Decoding Practical for Large-Scale, Long-Context RL Post-TrainingTL;DR : Rollout generation …