[Paper Review] Llama-Nemotron: Efficient Reasoning Models
Paper Link Hydragen: The Secret Weapon for Decoding Large Batches with Shared Prefixes up to 32× FasterTL;DRBy decomposing the prefix and …
25 min
Prefix Caching
Efficient Inference
Attention Optimization
FlashAttention
vLLM