The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
Paper Waking the MoE on SSD: Serving a 35B MoE at 20 tok/s on a 24 GB Desktop with Edge0One-line summary (TL;DR)Even at 4-bit, a 35B-class …
Paper Waking the MoE on SSD: Serving a 35B MoE at 20 tok/s on a 24 GB Desktop with Edge0One-line summary (TL;DR)Even at 4-bit, a 35B-class …
Paper Is “Lossless” Speculative Decoding Really Lossless? — The Numerical-Precision Trap Exposed by Reproducing OrthrusTL;DR — …
Paper ZGCM-1: How a 7B Dense Model Trades Blows with 235B Frontiers — A Fully Open, Ultra-Efficient Foundation Model Design ReportTL;DR — …
Paper SAS: A Simple Attention Sparsification That Learns Context Ranking Directly Without DistillationTL;DR — In post-training attention …
Paper CoVeR: Coverage-Based Pruning That Cuts Visual Tokens of Multi-View 3D Reasoning by up to 92% Using Geometry AloneOne-Line Summary …
Paper SQS: Fusing Pruning and Quantization into One Bayesian Learning — Where Spike-and-Slab Meets GMMTL;DR — Doing pruning and low-bit …
Linear Attention Rewritten with a Kalman Filter: Kalman Delta NetworksTL;DRThe fixed-size recurrent memory update of linear attention is …
Paper X-AuT: Pruning Audio Encoders in Speech LLMs with Behavioral Probes and Cross-Scale DistillationTL;DRX-AuT is a framework that …
Paper Deadline-Filling Prefill Chunks: SLOWeave — An Adaptive Chunking Scheduler for LLM Serving TL;DR — Instead of a fixed prefill chunk …
Paper Why Are Video LLMs Still So Expensive? — A Four-Stage Guide to Inference Efficiency MechanismsTL;DR — This paper classifies 125 …
Paper Link Hydragen: The Secret Weapon for Decoding Large Batches with Shared Prefixes up to 32× FasterTL;DRBy decomposing the prefix and …