The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
Paper Waking the MoE on SSD: Serving a 35B MoE at 20 tok/s on a 24 GB Desktop with Edge0One-line summary (TL;DR)Even at 4-bit, a 35B-class …
18 min
Mixture of Experts
Quantization
Efficient Inference
LLM Inference
Hardware Architecture
Open-Source LLM