abseil
0Korean renderings of Abseil's Tip of the Week — string_view, unique_ptr, temporaries and lifetimes, and the other C++ questions that keep coming back.
Korean renderings of Abseil's Tip of the Week — string_view, unique_ptr, temporaries and lifetimes, and the other C++ questions that keep coming back.
Reading an open-source inference engine end to end, pinned to one commit. Each part follows the execution path through a specific region of the code and the reasoning behind its design.
Paper reviews on making LLM inference faster and cheaper: speculative decoding, KV cache compression, prompt compression, sparsity, and heterogeneous pipelines.
Paper reviews on shrinking, evicting and reusing the KV cache: compression, quantization, eviction signals, offloading and prefix reuse.
Paper reviews on coping with long context: sparse attention, context compression and recompute, state-space models, and million-token serving.
Paper reviews on LLM serving and training across GPUs and nodes: prefill/decode disaggregation, pipeline parallelism, overlapping communication with compute, scheduling and energy.
Paper The KV Cache Is the New Memory Wall: Three Regimes and Five Domains of Long-Context LLM Inference TL;DR — Long-context LLM decode …
Paper WavePP — A Pipeline-Parallel Prefill Runtime That Burns In Prefix ReuseTL;DRBuilt on top of TensorRT-LLM, WavePP overlaps request …
Paper Phase Sensitivity: The ‘Periodic Weakness’ Created by Chunked KV-Cache Compression TL;DR — In models that use chunked …
Paper WorldAttention: System Co-Design of an Interactive Video World Model That Drops the Sliding WindowTL;DR — Text-controlled interactive …
Paper Who pays for the KV cache? — Kubernetes, gateway, and provider bills in one ledger TL;DR — The cost of a single AI feature is …