The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction 9/18/26 paper-review, mixture-of-experts, with-muse-spark-1-3-contributor-free 18 min
5. The decode path that produces one token — why separate kernels were made 9/17/26 code-series 16 min
6. Splitting the KV cache into 16-token blocks — reading non-contiguous blocks via the block table 9/17/26 code-series 18 min
7. Overlapping multiple requests with slots and a queue — implementing continuous batching 9/17/26 code-series 15 min
StepAudio 3 Gen Technical Report 9/17/26 paper-review, model-architecture, with-deepseek-v4-pro 5 min
How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus 9/16/26 paper-review, speculative-decoding, with-deepseek-v4-pro 13 min
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search 9/16/26 paper-review, long-context, with-deepseek-v4-pro 15 min
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking 9/15/26 paper-review, kv-cache, with-deepseek-v4-pro 20 min
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs 9/14/26 paper-review, sparsity-pruning, with-deepseek-v4-pro 4 min
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation 9/13/26 paper-review, model-architecture, with-deepseek-v4-pro 18 min
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions 9/13/26 paper-review, sparsity-pruning, with-deepseek-v4-pro 13 min
Kalman Delta Networks: Uncertainty-aware Associative Memory 9/12/26 paper-review, model-architecture, with-deepseek-v4-pro 16 min
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation 9/12/26 paper-review, sparsity-pruning, with-deepseek-v4-pro 11 min
Deadline-Aware Adaptive Prefill Chunking for Efficient Large Language Model Serving 9/11/26 paper-review, serving-systems, with-deepseek-v4-pro 15 min
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs 9/11/26 paper-review, long-context, with-deepseek-v4-pro 17 min
Miles v0.1: Production-Level Post-Training 9/10/26 paper-review, training-efficiency, with-deepseek-v4-pro 15 min
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training 9/10/26 paper-review, speculative-decoding, with-deepseek-v4-pro 11 min
Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference 9/9/26 paper-review, mixture-of-experts, with-muse-spark-1-3-contributor-free 20 min
VestigeKV: The NoPE-MLA KV Cache Carries Its Own Eviction Signal in a Vestigial Branch 9/9/26 paper-review, kv-cache, with-deepseek-v4-pro 15 min
ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding 9/8/26 paper-review, kv-cache, with-deepseek-v4-pro 17 min
Unlocking Lossless Speedups in LLMs via Discrete Diffusion 9/8/26 paper-review, speculative-decoding, with-deepseek-v4-pro 21 min
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference 9/7/26 paper-review, kv-cache, with-deepseek-v4-pro 14 min
Language Models Can Control Their Own Attention 9/7/26 paper-review, kv-cache, with-deepseek-v4-pro 14 min
Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving 9/7/26 paper-review, kv-cache, with-deepseek-v4-pro 16 min
SGD-KV: Summarization Guided KV Cache Compression 9/7/26 paper-review, kv-cache, with-deepseek-v4-pro 13 min
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers 9/7/26 paper-review, mixture-of-experts, with-deepseek-v4-pro 11 min
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM 9/7/26 paper-review, quantization, with-deepseek-v4-pro 14 min
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning 9/6/26 paper-review, kv-cache, with-qwen3-8-flash 21 min
[Paper Review] Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models 1/14/26 paper-review, with-gpt, llm-architecture, systems-ml 39 min
[Paper Review] Memory Retrieval and Consolidation in Large Language Models through Function Tokens 12/20/25 paper-review, interpretability, LLM Systems, with-gpt-5.2 23 min
[Paper Review] NVIDIA Nemotron 3: Efficient and Open Intelligence 12/16/25 paper-review, with-gpt-5.2 4 min
[Paper Review] Radial Attention: O(n log n) Sparse Attention with Energy Decay for Long Video Generation 12/15/25 paper-review, with-gpt-5.2 35 min
[Paper Review] Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin 10/23/25 paper-review, with-gpt 34 min
[Paper Review] Hardware-Efficient Attention for Fast Decoding 10/23/25 paper-review, with-gpt, Large Language Models, Model Efficiency, Compiler and Systems 35 min
[Paper Review] Towards Efficient and Practical GPU Multitasking in the Era of LLM 10/9/25 paper-review, with-gpt, LLM Systems, Model Serving and Scheduling, GPU Virtualization and Multi-Tenancy 36 min
[Paper Review] DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving 10/8/25 paper-review, LLM Serving and Systems, Model Optimization and Acceleration 15 min
[Paper Review] Marconi: Prefix Caching for the Era of Hybrid LLMs 10/8/25 paper-review, with-gpt, LLM Systems, Model Serving 11 min
[Paper Review] SGLang: Efficient Execution of Structured Language Model Programs 10/3/25 paper-review, with-gpt 39 min
[Paper Review] Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality 10/3/25 paper-review, with-gpt 19 min
[Paper Review] Inference-Time Hyper-Scaling with KV Cache Compression 7/29/25 paper-review, with-gpt 32 min
[Paper Review] Llama-Nemotron: Efficient Reasoning Models 7/29/25 paper-review, with-gpt, efficient-llm, system-optimization, inference-acceleration 25 min
[Paper Review] KIMI K2: OPEN AGENTIC INTELLIGENCE 7/26/25 paper-review, with-gpt, open-source, agentic-intelligence, RL-alignment, foundation-models 14 min
K-Beauty: Beyond 'Accidental Success' to 'Structural Growth' 7/26/25 Industry Analysis, Cosmetics Industry 4 min
[Paper Review] Peri-LN: Revisiting Normalization Layer in the Transformer Architecture 7/9/25 paper-review, with-gpt 24 min
[paper review] SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-bit Training 7/9/25 paper-review, with-gpt 23 min
[Paper Review] Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding 7/8/25 paper-review, with-gpt 17 min
DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition 7/8/25 paper-review, with-gpt 26 min
Code I/O: Condensing Reasoning Patterns via Code Input-Output Prediction 7/7/25 paper-review, with-gpt 36 min
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention 7/7/25 paper-review, with-gpt 36 min
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling 7/6/25 paper-review, with-gpt, Janus, DeepSeek 18 min
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation 7/2/25 paper-review, with-gpt 30 min