WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse 26. 10. 2. paper-review, serving-systems, with-deepseek-v4-pro 14분
Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression 26. 10. 1. paper-review, kv-cache, with-deepseek-v4-pro 16분
WorldAttention: An Efficient Attention Architecture for Interactive Video World Models 26. 9. 30. paper-review, kv-cache, with-deepseek-v4-pro 15분
Who Pays for the KV Cache? Attributing Shared AI Inference Spend Across Kubernetes and LLM Provider Bills 26. 9. 28. paper-review, serving-systems, with-deepseek-v4-pro 15분
SPECTRA: Adaptive Execution of Speculative Decoding on a Runtime-Reconfigurable Tiled Architecture 26. 9. 26. paper-review, hardware-accelerators, with-deepseek-v4-pro 13분
Adapting Tree-Structured Speculative Decoding to DeepSeek-V4 for Efficient Inference 26. 9. 25. paper-review, speculative-decoding, with-muse-spark-1-3-contributor-free 13분
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs 26. 9. 24. paper-review, kv-cache, with-muse-spark-1-3-contributor-free 4분
H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache 26. 9. 23. paper-review, speculative-decoding, with-muse-spark-1-3-contributor-free 14분
Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI Inference 26. 9. 22. paper-review, serving-systems, with-muse-spark-1-3-contributor-free 12분
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention 26. 9. 20. paper-review, attention-kernels, with-muse-spark-1-3-contributor-free 16분
Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches 26. 9. 18. paper-review, kv-cache, with-muse-spark-1-3-contributor-free 20분
The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction 26. 9. 18. paper-review, mixture-of-experts, with-muse-spark-1-3-contributor-free 16분
StepAudio 3 Gen Technical Report 26. 9. 17. paper-review, model-architecture, with-deepseek-v4-pro 6분
How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus 26. 9. 16. paper-review, speculative-decoding, with-deepseek-v4-pro 12분
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search 26. 9. 16. paper-review, long-context, with-deepseek-v4-pro 15분
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking 26. 9. 15. paper-review, kv-cache, with-deepseek-v4-pro 18분
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs 26. 9. 14. paper-review, sparsity-pruning, with-deepseek-v4-pro 5분
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation 26. 9. 13. paper-review, model-architecture, with-deepseek-v4-pro 17분
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions 26. 9. 13. paper-review, sparsity-pruning, with-deepseek-v4-pro 12분
Kalman Delta Networks: Uncertainty-aware Associative Memory 26. 9. 12. paper-review, model-architecture, with-deepseek-v4-pro 13분
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation 26. 9. 12. paper-review, sparsity-pruning, with-deepseek-v4-pro 12분
Deadline-Aware Adaptive Prefill Chunking for Efficient Large Language Model Serving 26. 9. 11. paper-review, serving-systems, with-deepseek-v4-pro 15분
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs 26. 9. 11. paper-review, long-context, with-deepseek-v4-pro 14분
Miles v0.1: Production-Level Post-Training 26. 9. 10. paper-review, training-efficiency, with-deepseek-v4-pro 14분
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training 26. 9. 10. paper-review, speculative-decoding, with-deepseek-v4-pro 11분
Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference 26. 9. 9. paper-review, mixture-of-experts, with-muse-spark-1-3-contributor-free 17분
VestigeKV: The NoPE-MLA KV Cache Carries Its Own Eviction Signal in a Vestigial Branch 26. 9. 9. paper-review, kv-cache, with-deepseek-v4-pro 14분
ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding 26. 9. 8. paper-review, kv-cache, with-deepseek-v4-pro 15분
Unlocking Lossless Speedups in LLMs via Discrete Diffusion 26. 9. 8. paper-review, speculative-decoding, with-deepseek-v4-pro 18분
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference 26. 9. 7. paper-review, kv-cache, with-deepseek-v4-pro 15분
Language Models Can Control Their Own Attention 26. 9. 7. paper-review, kv-cache, with-deepseek-v4-pro 14분
Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving 26. 9. 7. paper-review, kv-cache, with-deepseek-v4-pro 17분
SGD-KV: Summarization Guided KV Cache Compression 26. 9. 7. paper-review, kv-cache, with-deepseek-v4-pro 14분
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers 26. 9. 7. paper-review, mixture-of-experts, with-deepseek-v4-pro 11분
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM 26. 9. 7. paper-review, quantization, with-deepseek-v4-pro 14분
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning 26. 9. 6. paper-review, kv-cache, with-qwen3-8-flash 19분
[논문리뷰]: Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models 26. 1. 14. paper-review, with-gpt, llm-architecture, systems-ml 44분
[논문리뷰]: Memory Retrieval and Consolidation in Large Language Models through Function Tokens 25. 12. 20. paper-review, interpretability, LLM Systems, with-gpt-5.2 51분
[논문리뷰]: Radial Attention: O(n log n) Sparse Attention with Energy Decay for Long Video Generation 25. 12. 15. paper-review, with-gpt-5.2 37분
[논문 리뷰] Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin 25. 10. 23. paper-review, with-gpt 36분
[논문 리뷰] Hardware-Efficient Attention for Fast Decoding 25. 10. 23. paper-review, with-gpt, Large Language Models, Model Efficiency, Compiler and Systems 35분
[논문리뷰]Towards Efficient and Practical GPU Multitasking in the Era of LLM 25. 10. 9. paper-review, with-gpt, LLM Systems, Model Serving and Scheduling, GPU Virtualization and Multi-Tenancy 41분
[논문리뷰] DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving 25. 10. 8. paper-review, LLM Serving and Systems, Model Optimization and Acceleration 40분
[논문리뷰] Marconi: Prefix Caching for the Era of Hybrid LLMs 25. 10. 8. paper-review, with-gpt, LLM Systems, Model Serving 41분
[논문리뷰] SGLang: Efficient Execution of Structured Language Model Programs 25. 10. 3. paper-review, with-gpt 44분
[논문리뷰] Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality 25. 10. 3. paper-review, with-gpt 44분
[논문리뷰] Llama-Nemotron: Efficient Reasoning Models 25. 7. 29. paper-review, with-gpt, efficient-llm, system-optimization, inference-acceleration 29분
[논문리뷰] KIMI K2: OPEN AGENTIC INTELLIGENCE 25. 7. 26. paper-review, with-gpt, open-source, agentic-intelligence, RL-alignment, foundation-models 31분
[논문리뷰] Peri-LN: Revisiting Normalization Layer in the Transformer Architecture 25. 7. 9. paper-review, with-gpt 26분
[논문리뷰] SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-bit Training 25. 7. 9. paper-review, with-gpt 30분
[논문리뷰] Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding 25. 7. 8. paper-review, with-gpt 35분
DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition 25. 7. 8. paper-review, with-gpt 26분
Code I/O: Condensing Reasoning Patterns via Code Input-Output Prediction 25. 7. 7. paper-review, with-gpt 34분
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention 25. 7. 7. paper-review, with-gpt 34분
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling 25. 7. 6. paper-review, with-gpt, Janus, DeepSeek 34분
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation 25. 7. 2. paper-review, with-gpt 29분
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts 25. 7. 1. paper-review, with-gpt 31분
Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning 25. 7. 1. paper-review, with-gpt 31분
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence 25. 6. 30. paper-review, with-gpt 31분
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models 25. 6. 29. paper-review, with-gpt 40분
DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior 25. 6. 29. paper-review, with-gpt, 3D, Diffusion 32분
Accelerated Test-Time Scaling with Model-Free Speculative Sampling 25. 6. 26. paper-review, with-gpt-o3 30분
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction 25. 6. 26. paper-review, with-gpt-o3 27분
Compress, Gather, and Recompute: REFORMingLong-Context Processing in Transformers 25. 6. 24. paper-review, with-gpt-o3 28분
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities 25. 6. 23. paper-review, with-gpt-o3 34분
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching 25. 6. 19. paper-review, with-gemini-2.5-pro-preview 39분
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention 25. 6. 19. paper-review, with-gemini-2.5-pro-preview 38분
MMInference: Accelerating Pre-filling for Long-Context Visual Language Models via Modality-Aware Permutation Sparse Attention 25. 6. 19. paper-review, with-gemini-2.5-pro-preview 43분
PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters 25. 6. 19. paper-review, with-gemini-2.5-pro-preview 42분
Slim attention: cut your context memory in half without loss– K-cache is all you need for MHA 25. 6. 16. paper-review, with-gemini-2.5-pro-preview 30분
Towards Economical Inference: Enabling DeepSeek’s Multi-Head Latent Attention in Any Transformer-based LLMs 25. 6. 16. paper-review, with-gemini-2.5-pro-preview 36분
TransMLA: Multi-Head Latent Attention Is All You Need 25. 6. 16. paper-review, with-gemini-2.5-pro-preview 34분
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression 25. 6. 16. paper-review, with-gemini-2.5-pro-preview 32분
A Bring-Your-Own-Model Approach for ML-Driven Storage Placement in Warehouse-Scale Computers 25. 6. 10. paper-review, with-gemini-2.5-pro-preview, MLSYS2025 33분
Know Where You're Uncertain When Planning with Multimodal Foundation Models: A Formal Framework 25. 6. 10. paper-review, with-gemini-2.5-pro-preview, MLSYS2025 35분
ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation 25. 6. 10. paper-review, with-gemini-2.5-pro-preview, MLSYS2025 30분
Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling 25. 6. 10. paper-review, with-gemini-2.5-pro-preview, MLSYS2025 27분
Supply-Chain Attacks in Machine Learning Frameworks 25. 6. 10. paper-review, with-gemini-2.5-pro-preview, MLSYS2025 28분
Accelerating MoE Model Inference with Expert Sharding 25. 6. 5. paper-review, with-gemini-2.5-pro-preview 34분
FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference 25. 6. 5. paper-review, with-gemini-2.5-pro-preview, MLSYS2025 35분
ScaleFusion: Scalable Inference of Spatial-Temporal Diffusion Transformers for High-Resolution Long Video Generation 25. 6. 5. paper-review, with-gemini-2.5-pro-preview, MLSYS2025 41분
SOLA: Optimizing SLO Attainment for Large Language Model Serving with State-Aware Scheduling 25. 6. 5. paper-review, with-gemini-2.5-pro-preview, MLSYS2025 38분
XGRAMMAR: FLEXIBLE AND EFFICIENT STRUCTURED GENERATION ENGINE FOR LARGE LANGUAGE MODELS 25. 6. 2. paper-review, with-gemini-2.5-pro-preview, MLSYS2025 28분
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures 25. 5. 17. paper-review, with-gemini-2.5-pro-preview 70분
RODIMUS*: BREAKING THE ACCURACY-EFFICIENCY TRADE-OFF WITH EFFICIENT ATTENTIONS 25. 5. 17. paper-review, with-gemini-2.5-pro-preview 46분
Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model 25. 4. 16. paper-review, with-gpt 16분
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts 25. 4. 14. paper-review, with-gpt 24분
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism 25. 4. 14. paper-review, with-gpt 22분
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention 25. 4. 14. paper-review, with-gpt 24분
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching 25. 4. 13. paper-review, with-gpt 22분
FLEX ATTENTION: A PROGRAMMING MODEL FOR GENERATING OPTIMIZED ATTENTION KERNELS 25. 4. 7. paper-review, with-gpt 29분
LeanAttention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers 25. 4. 7. paper-review, with-gpt 27분
AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds 25. 4. 2. paper-review, with-gpt 16분
SparseTransX: Efficient Training of Translation-Based Knowledge Graph Embeddings Using Sparse Matrix Operations 25. 4. 2. paper-review, with-gpt 16분
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference 25. 3. 31. paper-review, with-gpt, MLSYS2025 25분
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training 25. 3. 25. paper-review, with-gpt 25분
SELF-DATA DISTILLATION FOR RECOVERING QUALITY IN PRUNED LARGE LANGUAGE MODELS 25. 3. 25. paper-review, with-gpt 21분
On Distributed Larger-Than-Memory Subset Selection With Pairwise Submodular Functions 25. 3. 24. paper-review, with-gpt 23분
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention 25. 3. 24. paper-review, with-gpt 13분
TRAINING ULTRA LONG CONTEXT LANGUAGE MODEL WITH FULLY PIPELINED DISTRIBUTED TRANSFORMER 25. 3. 24. paper-review, with-gpt 21분
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving 25. 3. 18. paper-review, with-gpt 30분
AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution 25. 3. 17. paper-review, with-gpt, MLSYS2025 19분
DIFFSERVE: EFFICIENTLY SERVING TEXT-TO-IMAGE DIFFUSION MODELS WITH QUERY-AWARE MODEL SCALING 25. 3. 17. paper-review, with-gpt, MLSYS2025 31분
EFFICIENT LLM INFERENCE USING DYNAMIC INPUT PRUNING AND CACHE-AWARE MASKING 25. 3. 12. paper-review, with-gpt, MLSYS2025 38분
LAVA: LIFETIME-AWARE VM ALLOCATION WITH LEARNED DISTRIBUTIONS AND ADAPTATION TO MISPREDICTIONS 25. 3. 11. paper-review, with-gpt, MLSYS2025 40분
TurboAttention: Efficient Attention Approximation for High Throughputs LLMs 25. 3. 11. paper-review, with-gpt, MLSYS2025 35분
A PRACTICAL CROSS-LAYER APPROACH FOR ML-DRIVEN STORAGE PLACEMENT IN WAREHOUSE-SCALE COMPUTERS 25. 3. 10. paper-review, with-gpt, MLSYS2025 38분
Scaling Deep Learning Training with MPMD Pipeline Parallelism 25. 3. 10. paper-review, with-gpt, MLSYS2025 29분
LSERVE: EFFICIENT LONG-SEQUENCE LLM SERVING WITH UNIFIED SPARSE ATTENTION 25. 3. 6. paper-review, with-gpt, MLSYS2025 30분
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments 25. 3. 6. paper-review, with-gpt, MLSYS2025 24분
VOLUT: EFFICIENT VOLUMETRIC STREAMING ENHANCED BY LUT-BASED SUPER-RESOLUTION 25. 3. 6. paper-review, with-gpt, MLSYS2025 28분
Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences 25. 3. 4. paper-review, with-gpt 20분
Forget the Data and Fine-Tuning! Just Fold the Network to Compress 25. 3. 4. paper-review, with-gpt 33분
HEXGEN-2: DISAGGREGATED GENERATIVE INFERENCE OF LLMS IN HETEROGENEOUS ENVIRONMENT 25. 2. 25. paper-review, with-gpt, ICLR2025 39분
Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding 25. 2. 25. paper-review, with-gpt, ICLR2025 39분
You OnlyPruneOnce: DESIGNING CALIBRATION-FREE MODEL COMPRESSION WITH POLICY LEARNING 25. 2. 25. paper-review, with-gpt, ICLR2025 36분
FlashMask: Efficient and Rich Mask Extension of FlashAttention 25. 2. 24. paper-review, with-gpt, ICLR2025 39분
TypedThinker: Typed Thinking Improves Large Language Model Reasoning 25. 2. 24. paper-review, with-gpt, ICLR2025 44분
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid 25. 2. 17. paper-review, with-gpt 27분
Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign 25. 2. 13. paper-review, with-gpt 35분
SmolLM2: When Smol Goes Big Data-Centric Training of a Small Language Mode 25. 2. 13. paper-review, with-gpt 39분
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning 25. 2. 12. paper-review, with-gpt 37분
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding 25. 2. 12. paper-review, with-gpt 39분
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence 25. 2. 11. paper-review, with-gpt 40분
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at AnyResolution 25. 2. 11. paper-review, with-gpt, Qwen 41분
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search 25. 2. 10. paper-review, with-gpt, DeepSeek 45분
Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models 25. 2. 10. paper-review, with-gpt 37분
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data 25. 2. 9. paper-review, with-gpt, DeepSeek 39분
DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence 25. 2. 7. paper-review, with-gpt 37분
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models 25. 2. 7. paper-review, with-gpt 38분
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism 25. 2. 5. paper-review, with-gpt 37분
Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models 25. 2. 5. paper-review, with-gpt 40분
Qwen-VL: AVersatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond 25. 2. 4. paper-review, with-gpt 34분
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation 25. 2. 3. paper-review, with-gpt 40분
A Hardware Evaluation Framework for Large Language Model Inference 25. 1. 21. paper-review, with-gpt 25분
Compressed Context Memory For Online Language Model Interaction 25. 1. 21. paper-review, with-gpt 28분
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs 25. 1. 20. paper-review, with-gpt 23분
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs 25. 1. 20. paper-review, with-gpt 29분
TAIPAN: EFFICIENT AND EXPRESSIVE STATE SPACE LANGUAGE MODELS WITH SELECTIVE ATTENTION 25. 1. 20. paper-review, with-gpt 35분
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication 25. 1. 20. paper-review, with-gpt 21분
SANA: EFFICIENT HIGH-RESOLUTION IMAGE SYN THESIS WITH LINEAR DIFFUSION TRANSFORMERS 25. 1. 15. paper-review, with-gpt 26분
Block Transformer: Global-to-Local Language Modeling for Fast Inference 25. 1. 15. paper-review, with-gpt 30분
MobiLlama: Towards Accurate and Lightweight Fully Transparent GPT 25. 1. 15. paper-review, with-gpt 23분
Rethinking Optimization and Architecture for Tiny Language Models 25. 1. 15. paper-review, with-gpt 11분
Distributed Inference and Fine-tuning of Large Language Models Over The Internet 25. 1. 2. paper-review, with-gpt 23분
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism 25. 1. 2. paper-review, with-gpt 21분
Gated Linear Attention Transformers with Hardware-Efficient Training 25. 1. 2. paper-review, with-gpt 28분
LLM in a flash : Efficient Large Language Model Inference with Limited Memory 25. 1. 2. paper-review, with-gpt 25분
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge 24. 12. 31. paper-review, with-gpt 24분
Improving alignment of dialogue agents via targeted human judgements 24. 12. 31. paper-review, with-gpt 27분
SCCA: Shifted Cross Chunk Attention for long contextual semantic expansion 24. 12. 30. paper-review, with-gpt 22분
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection 24. 12. 26. paper-review, with-gpt 22분
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context 24. 12. 26. paper-review, with-gpt 28분
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding 24. 12. 26. paper-review, with-gpt 22분
The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction 24. 12. 26. paper-review, with-gpt 26분
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving 24. 12. 24. paper-review, with-gpt 25분
SAGEATTENTION: ACCURATE 8-BIT ATTENTION FOR PLUG-AND-PLAY INFERENCE ACCELERATION 24. 12. 24. paper-review, with-gpt 24분
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality 24. 12. 24. paper-review, with-gpt 8분
FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs 24. 12. 23. paper-review, with-gpt 23분
SimpleFSDP: Simpler Fully Sharded Data Parallel with torch.compile 24. 12. 23. paper-review, with-gpt 22분
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression 24. 12. 20. paper-review, with-gpt 25분
Efficient LLM Inference with I/O-Aware Partial KV Cache Recomputation 24. 12. 20. paper-review, with-gpt 24분
Large Concept Models: Language Modeling in a Sentence Representation Space 24. 12. 20. paper-review, with-gpt 26분
SageAttention2 Technical Report: Accurate 4 Bit Attention for Plug-and-play Inference Acceleration 24. 12. 20. paper-review, with-gpt 25분
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference 24. 12. 20. paper-review, with-gpt 21분
Efficient Memory Management for Large Language Model Serving with PagedAttention 24. 12. 19. paper-review, with-gpt 28분
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding 24. 12. 19. paper-review, with-gpt 30분
GSPMD: General and Scalable Parallelization for ML Computation Graphs 24. 12. 19. paper-review, with-gpt 24분
Orca: Progressive Learning from Complex Explanation Traces of GPT-4 24. 12. 19. paper-review, with-gpt 29분
Fast Inference of Mixture-of-Experts Language Models with Offloading 24. 12. 18. paper-review, with-gpt 25분
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning 24. 12. 18. paper-review, with-gpt 29분
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision 24. 12. 18. paper-review, with-gpt 26분
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness 24. 12. 18. paper-review, with-gpt 26분
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache 24. 12. 18. paper-review, with-gpt 26분
SuperServe: Fine-Grained Inference Serving for Unpredictable Workloads 24. 12. 18. paper-review, with-gpt 9분
Fast and Effective Weight Update for Pruned Large Language Models 24. 12. 17. paper-review, with-gpt 25분
FFSplit: Split Feed-Forward Network For Optimizing Accuracy-Efficiency Trade-off in Language Model Inference 24. 12. 17. paper-review, with-gpt 23분
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs 24. 12. 17. paper-review, with-gpt 27분
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models 24. 12. 17. paper-review, with-gpt 21분
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts 24. 12. 17. paper-review, with-gpt 27분
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference 24. 12. 16. paper-review, with-gpt 30분
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving 24. 12. 16. paper-review, with-gpt 36분
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads 24. 12. 16. paper-review, with-gpt 28분
INFERFLOW: AN EFFICIENT AND HIGHLY CONFIG URABLE INFERENCE ENGINE FOR LARGE LANGUAGE MODELS 24. 12. 16. paper-review, with-gpt 30분
MEDUSA: Simple LLMInference Acceleration Framework with Multiple Decoding Heads 24. 12. 16. paper-review, with-gpt 24분
Break the Sequential Dependency of LLM Inference Using LOOKAHEAD DECODING 24. 12. 15. paper-review, with-gpt 34분
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design 24. 12. 15. paper-review, with-gpt 51분
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture 24. 12. 14. paper-review, with-gpt 42분
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inferenc 24. 12. 14. paper-review, with-gpt 17분
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference 24. 12. 14. paper-review, with-gpt 41분
RelayAttention for Efficient Large Language Model Serving with Long System Prompts 24. 12. 14. paper-review, with-gpt 53분
TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models 24. 12. 14. paper-review, with-gpt 13분
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers 24. 12. 13. paper-review, with-gpt 15분
GQA:Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints 24. 12. 13. paper-review, with-gpt 17분
H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Model 24. 12. 13. paper-review, with-gpt 17분
Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time 24. 12. 13. paper-review, with-gpt 16분
FlexLLM: A System for Co-Serving Large Language Model Inference and Parameter-Efficient Finetuning 24. 12. 12. paper-review, with-gpt 15분
LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models 24. 12. 12. paper-review, with-gpt 19분
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs 24. 12. 12. paper-review, with-gpt 14분
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization 24. 12. 12. paper-review, with-gpt 15분
Benchmarks as Limits to Arbitrage: Understanding the Low-Volatility Anomaly 24. 12. 10. paper-review, with-gpt, finance 13분
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model 24. 12. 10. paper-review, with-gpt, LLM 19분
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving 24. 12. 10. paper-review, with-gpt, LLM-Inference 21분
CHAI: Clustered Head Attention for Efficient LLM Inference 24. 12. 10. paper-review, with-gpt, LLM-Inference 19분
High Idiosyncratic Volatility and Low Returns: International and Further U.S. Evidence 24. 12. 10. paper-review, with-gpt, finance 12분
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression 24. 12. 10. paper-review, with-gpt, LLM-Inference 2분
DeepCache: Accelerating Diffusion Models for Free 24. 12. 9. paper-review, with-gpt, LLM-Inference 19분
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference 24. 12. 9. paper-review, with-gpt, LLM-Inference 19분
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache 24. 12. 9. paper-review, with-gpt, LLM-Inference 15분
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization 24. 12. 9. paper-review, with-gpt, LLM-Inference 14분
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More 24. 12. 9. paper-review, with-gpt, LLM-Inference 17분
Improving Language Understanding by Generative Pre-Training 24. 12. 8. paper-review, with-gpt, LLM 19분
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU 24. 12. 8. paper-review, with-gpt, LLM-Inference 14분
QAQ: Quality Adaptive Quantization for LLM KV Cache 24. 12. 8. paper-review, with-gpt, LLM-Inference 17분
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching 24. 12. 6. paper-review, with-gpt, LLM-Inference 16분
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 24. 12. 6. paper-review, with-gpt, LLM 19분
Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference 24. 12. 6. paper-review, with-gpt, Dynamic Memory Compression, LLM-Inference 14분
Fast Inference from Transformers via Speculative Decoding 24. 12. 6. paper-review, with-gpt, LLM-Inference, speculative-decoding 3분
FASTDECODE: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines 24. 12. 6. paper-review, with-gpt, FASTDECODE, LLM-Inference 15분
LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression 24. 12. 6. paper-review, with-gpt, LLMLingua-2, LLM-Inference 15분
Direct Preference Optimization: Your Language Model is Secretly a Reward Model 24. 12. 5. paper-review, with-gpt 13분
HIERARCHICAL CONTEXT MERGING: BETTER LONG CONTEXT UNDERSTANDING FOR PRE-TRAINED LLMS 24. 12. 5. paper-review, with-gpt 12분
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving 24. 12. 5. paper-review, with-gpt 14분
The Flan Collection: Designing Data and Methods for Effective Instruction Tuning 24. 12. 5. paper-review, with-gpt 24분
Transformer-Lite: High-efficiency Deployment of Large Language Models on Mobile Phone GPUs 24. 12. 5. paper-review, with-gpt 20분
Tree of Thoughts: Deliberate Problem Solving with Large Language Models 24. 12. 5. paper-review, with-gpt 23분
CORM: Cache Optimization with Recent Message for Large Language Model Inference 24. 12. 4. paper-review, with-gpt 18분
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention 24. 12. 4. paper-review, with-gpt 22분
Retrieval Head Mechanistically Explains Long-Context Factuality 24. 12. 4. paper-review, with-gpt 23분
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget 24. 12. 4. paper-review, with-gpt 24분
Toward Inference-optimal Mixture-of-Expert Large Language Models 24. 12. 4. paper-review, with-gpt 14분
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 24. 12. 3. paper-review, with-gpt 19분
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases 24. 12. 3. paper-review, with-gpt 7분
Parallel Decoding via Hidden Transfer for Lossless Large Language Model Acceleration 24. 12. 3. paper-review, with-gpt 22분
PowerInfer-2: Fast Large Language Model Inference on a Smartphone 24. 12. 3. paper-review, with-gpt 28분
RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation 24. 12. 3. paper-review, with-gpt 16분
Layer-Condensed KV Cache for Efficient Inference of Large Language Models 24. 12. 2. paper-review, with-gpt 16분
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference 24. 12. 2. paper-review, with-gpt 17분
SKVQ:Sliding-window Key and Value Cache Quantization for Large Language Models 24. 12. 2. paper-review, with-gpt 23분
You Only Cache Once: Decoder-Decoder Architectures for Language Models 24. 12. 2. paper-review, with-gpt 16분
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models 24. 11. 29. paper-review, with-gpt 19분
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling 24. 11. 29. paper-review, with-gpt 18분
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention 24. 11. 29. paper-review, with-gpt 22분
ASimple and Effective L2 Norm-Based Strategy for KV Cache Compression 24. 11. 28. paper-review, with-gpt 14분
Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters 24. 11. 28. paper-review, with-gpt 18분
CItruS : ChunkedInstruction-aware State Eviction for Long Sequence Modeling 24. 11. 28. paper-review, with-gpt 15분
MLKV:Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding 24. 11. 28. paper-review, with-gpt 15분
Dynamic Discriminative Operations (D2O) for Efficient Generative Inference of Large Language Models 24. 11. 27. paper-review, with-gpt 15분
LOOK-M:Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference 24. 11. 27. paper-review, with-gpt 16분
MODEL TELLS YOU WHERE TO MERGE: ADAPTIVE KV CACHE MERGING FOR LLMS ON LONG-CONTEXT TASKS 24. 11. 27. paper-review, with-gpt 19분
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference 24. 11. 26. paper-review, with-gpt 22분
Keep the Cost Down: A Review on Methods to Optimize LLM’s KV Cache Consumption. 24. 11. 26. paper-review, with-gpt 6분
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference 24. 11. 26. paper-review, with-gpt 22분
PQCache: Product Quantization-based KVCache for Long Context LLM Inference 24. 11. 26. paper-review, with-gpt 21분
NACL: AGeneral and Effective KV Cache Eviction Framework for LLMs at Inference Time 24. 11. 25. paper-review, with-gpt 16분
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval 24. 11. 25. paper-review, with-gpt 18분
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction 24. 11. 21. paper-review, with-gpt 28분
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs 24. 11. 21. paper-review, with-gpt 18분
KV-COMPRESS: Paged KV-Cache Compression with Variable Compression Rates per Attention Head 24. 11. 21. paper-review, with-gpt 26분
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads 24. 11. 21. paper-review, with-gpt 23분
TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning 24. 11. 21. paper-review, with-gpt 23분
DUOATTENTION: EFFICIENT LONG-CONTEXT LLM INFERENCE WITH RETRIEVAL AND STREAMING HEADS 24. 11. 20. paper-review, with-gpt 20분
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy 24. 11. 20. paper-review, with-gpt 15분
SPARSEVLM: VISUAL TOKEN SPARSIFICATION FOR EFFICIENT VISION-LANGUAGE MODEL INFERENCE 24. 11. 20. paper-review, with-gpt 19분
SWIFTKV: FAST PREFILL-OPTIMIZED INFERENCE WITH KNOWLEDGE-PRESERVING MODEL TRANSFORMATION 24. 11. 20. paper-review, with-gpt 19분
TIDALDECODE: FAST AND ACCURATE LLM DECOD ING WITH POSITION PERSISTENT SPARSE ATTENTION 24. 11. 20. paper-review, with-gpt 16분
TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training 24. 11. 20. paper-review, with-gpt 17분
A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference 24. 11. 19. paper-review, with-gpt 19분
SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction 24. 11. 19. paper-review, with-gpt 19분
Recycled Attention: Efficient inference for long-context language models 24. 11. 18. paper-review, with-gpt 22분
Squeezed Attention: Accelerating Long Context Length LLM Inference 24. 11. 18. paper-review, with-gpt 24분
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection 24. 11. 18. paper-review, with-gpt 35분
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration 24. 11. 18. paper-review, with-gpt 29분
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models 24. 11. 14. paper-review, with-gpt 16분
Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models 24. 11. 14. paper-review, with-gpt 22분
HART Efficient Visual Generation with Hybrid Autoregressive Transformer 24. 11. 14. paper-review, with-gpt 21분
Learning Transferable Visual Models From Natural Language Supervision 24. 11. 14. paper-review, with-gpt 27분
The CoT Collection Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning 24. 11. 14. paper-review, with-gpt 22분
Condition-Aware Neural Network for Controlled Image Generation 24. 11. 13. paper-review, with-gpt 24분
DistriFusion Distributed Parallel Inference for High-Resolution Diffusion Models 24. 11. 13. paper-review, with-gpt 33분
FastComposer Tuning-Free Multi-Subject Image Generation with Localized Attention 24. 11. 13. paper-review, with-gpt 25분
VILA-U a Unified Foundation Model Integrating Visual Understanding and Generation 24. 11. 13. paper-review, with-gpt 23분
Batch Calibration Rethinking Calibration for In-Context Learning and Prompt Engineering 24. 11. 12. paper-review, with-gpt 20분
LiteMoE Customizing On-device LLM Serving via Proxy Submodel Tuning 24. 11. 12. paper-review, with-gpt 38분
ShadowKV KV Cache in Shadows for High-Throughput Long-Context LLM Inference 24. 11. 12. paper-review, with-gpt 18분
EPIC Efficient Position-Independent Context Caching for Serving Large Language Models 24. 11. 11. paper-review, with-gpt 11분
RAG4ITOps A Supervised Fine-Tunable and Comprehensive RAG Framework for IT Operations and Maintenance 24. 11. 11. paper-review, with-gpt 15분
Scientific Beta Multi-Beta Multi-Strategy Indices Implementing Multi-Factor Equity Portfolios with Smart Factor Indices 24. 11. 11. paper-review, with-gpt, finance 14분
ALPINE Unveiling the Planning Capability of Autoregressive Learning in Language Models 24. 11. 10. paper-review, with-gpt 10분
Capital asset prices A theory of market equilibrium under conditions of risk 24. 11. 10. paper-review, with-gpt, finance 10분
DynamoLLM Designing LLM Inference Clusters for Performance and Energy Efficiency 24. 11. 10. paper-review, with-gpt 12분
HYSYNTH Context-Free LLM Approximation for Guiding Program Synthesis 24. 11. 10. paper-review, with-gpt 11분
MInference 1.0 Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention 24. 11. 10. paper-review, with-gpt 23분
Enabling Tensor Language Model to Assist in Generating High-Performance Tensor Programs for Deep Learning 24. 11. 7. paper-review, with-gpt 3분
Meta Large Language Model Compiler Foundation Models of Compiler Optimization 24. 11. 7. paper-review, with-gpt 4분
BUZZ Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference 24. 11. 6. paper-review, with-gpt 5분
CDMPP:ADevice-Model Agnostic Framework for Latency Prediction of Tensor Programs 24. 11. 6. paper-review, with-gpt 6분
Don't Look Twice Faster Video Transformers with Run-Length Tokenization 24. 11. 6. paper-review, with-gpt 3분
FLUX Fast Software-based Communication Overlap On GPUs Through Kernel Fusion 24. 11. 6. paper-review, with-gpt 7분
KVSharer Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing 24. 11. 6. paper-review, with-gpt 7분
SpotServe Serving Generative Large Language Models on Preemptible Instances 24. 11. 5. paper-review, with-gpt 13분
GraphPipe Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism 24. 11. 4. paper-review, with-gpt 15분
Helix Distributed Serving of Large Language Models via Max-Flow on Heterogeneous GPUs 24. 11. 4. paper-review, with-gpt 4분
Sequoia Scalable, Robust, and Hardware-aware Speculative Decoding 24. 11. 4. paper-review, with-gpt 8분
SpecExec Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices 24. 11. 4. paper-review, with-gpt 7분
Breaking the Curse of Quality Saturation with User-Centric Ranking 24. 11. 3. paper-review, with-gpt 3분
Hard Disk Drive Failure Analysis and Prediction: An Industry View 24. 11. 3. paper-review, with-gpt 4분
MEGABYTE Predicting Million-byte Sequences with Multiscale Transformers 24. 11. 3. paper-review, with-gpt 10분
Reasoning over Public and Private Data in Retrieval-Based Systems 24. 11. 3. paper-review, with-gpt 3분
FlexGen High-Throughput Generative Inference of Large Language Models with a Single GPU 24. 11. 1. paper-review, with-gpt 17분
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches 24. 11. 1. paper-review, with-gpt 7분
Quest Query-Aware Sparsity for Efficient Long-Context LLM Inference 24. 11. 1. paper-review, with-gpt 9분
What Matters in Transformers? Not All Attention is Needed Fusion 24. 11. 1. paper-review, with-gpt 5분
Better & Faster Large Language Models via Multi-token Prediction 24. 10. 31. paper-review, with-gpt 14분
CacheBlend Fast Large Language Model Serving for RAG with Cached Knowledge Fusion 24. 10. 31. paper-review, with-gpt 9분
Keyformer KV Cache Reduction through Key Tokens Selection for Efficient Generative Inference 24. 10. 31. paper-review, with-gpt 16분
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve 24. 10. 31. paper-review, with-gpt 19분
간단논문 정리 End-to-End Deep Learning of Optimization Heuristics (PACT 17) 21. 2. 12. compiler, ML, paper-review 1분
간단논문 정리 Fast and Effective Orchestration of Compiler Optimizations(Zhelong Pan,Rudolf Eigenmann;Purdue University ;CGO’06) 21. 2. 12. compiler, ML, paper-review 1분
간단논문 정리 TVM An Automated End-to-End Optimizing Compiler for Deep Learning (OSDI 18) 21. 2. 12. compiler, ML, paper-review 1분
논문 정리 Chameleon Adaptive Code Optimization for Expedited Deep Neural Network Compilation(ICLR 2020) 21. 2. 12. compiler, ML, paper-review 2분
논문 정리 LLVM A Compilation Framework for Lifelong Program Analysis & Transformation(CGO 04) 21. 2. 12. compiler, paper-review 2분
논문 정리 NeuroVectorizer End-to-End Vectorization with Deep Reinforcement Learning (CGO 20) 21. 2. 12. compiler, ML, paper-review 1분