LLM-Inference

Paper reviews about running a trained model faster and cheaper. The axes that keep coming back:

  • Drafting tokens ahead and verifying them — speculative decoding
  • Shrinking the KV cache — quantization, sparsity-aware eviction, dynamic memory compression
  • Shrinking the input side — prompt compression
  • Splitting the hardware — heterogeneous pipelines, serving on consumer-grade GPUs
0 pages

No content

Back to home