The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
Paper Waking the MoE on SSD: Serving a 35B MoE at 20 tok/s on a 24 GB Desktop with Edge0One-line summary (TL;DR)Even at 4-bit, a 35B-class …
Paper Waking the MoE on SSD: Serving a 35B MoE at 20 tok/s on a 24 GB Desktop with Edge0One-line summary (TL;DR)Even at 4-bit, a 35B-class …
Paper Is “Lossless” Speculative Decoding Really Lossless? — The Numerical-Precision Trap Exposed by Reproducing OrthrusTL;DR — …
Paper CoVeR: Coverage-Based Pruning That Cuts Visual Tokens of Multi-View 3D Reasoning by up to 92% Using Geometry AloneOne-Line Summary …
Paper Deadline-Filling Prefill Chunks: SLOWeave — An Adaptive Chunking Scheduler for LLM Serving TL;DR — Instead of a fixed prefill chunk …
Paper “Let’s put an OS on the GPU”: A proposal for a GPU multitasking OS layer for the LLM eraOne-line summary (TL;DR)In …
![[Paper Review] SGLang: Efficient Execution of Structured Language Model Programs](https://cdn.bytez.com/mobilePapers/v2/neurips/94872/images/20-0.png)
Paper Link SGLang & RadixAttention: How Execution Optimization for “LM Programs” Achieved a 6.4x SpeedupTL;DRBy combining a …