Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving
Paper Same Request, Different Answer: Prefix Caching Makes Serving Non-Reproducible, and Quantization Amplifies It TL;DR — Prefix caching is …
16 min
Prefix Caching
Software Engineering
![[Paper Review] Marconi: Prefix Caching for the Era of Hybrid LLMs](https://pbs.twimg.com/media/GdyLXO9W4AADox0.jpg)