The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
Paper Waking the MoE on SSD: Serving a 35B MoE at 20 tok/s on a 24 GB Desktop with Edge0One-line summary (TL;DR)Even at 4-bit, a 35B-class …
Paper Waking the MoE on SSD: Serving a 35B MoE at 20 tok/s on a 24 GB Desktop with Edge0One-line summary (TL;DR)Even at 4-bit, a 35B-class …
Paper Miles v0.1: A “verified, clean, and scalable” full-stack system for frontier RL post-trainingTL;DR: Miles is a full-stack …
Paper Cache-Learning Routers: Cache-Aware Joint Router Adaptation for Memory-Constrained MoE InferenceTL;DRStarting from the observation …
Paper SMELT: Scaling Laws for Compute-Matched MoE Loop TransformersTL;DRLoop transformers scale depth by executing a layer block repeatedly, …
![[Paper Review] Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models](https://paper-assets.alphaxiv.org/figures/2601.07372v1/img-0.jpeg)
Paper Engram: The Second Sparsity Axis After Conditional Computation (MoE), Conditional MemoryOne-Line Summary (TL;DR)Engram adds a …
![[Paper Review] NVIDIA Nemotron 3: Efficient and Open Intelligence](https://developer-blogs.nvidia.com/wp-content/uploads/2025/12/image3-8-png.webp)
Paper NVIDIA Nemotron 3: Pushing the “Accuracy/Throughput” Frontier with a Hybrid Mamba–Transformer MoENemotron 3 combines an …
Paper Link Qwen 3: The Evolution of a Giant MoE Language Model with Adjustable Reasoning DepthTL;DR (in one line)Qwen 3 couples a …
![[Paper Review] Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding](https://www.storagereview.com/wp-content/uploads/2025/07/image2-2-png-e1752234784623.webp)
Paper Link Helix Parallelism: Breaking the Latency-Throughput Wall of Ultra-Long LLM DecodingTL;DRHelix Parallelism schedules Attention and …