Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI Inference
Paper How Many Prefills and How Many Decodes Within a Power Budget: The Capacity-Power Pareto Front for PD-Disaggregated InferenceOne-Line …
14 min
LLM Inference
Efficient Inference
Distributed Computing
KV Cache
Performance
GPU Acceleration