<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Distributed Computing on Jaehun's Blog</title><link>https://jaehun.me/tags/distributed-computing/</link><description>Recent content in Distributed Computing on Jaehun's Blog</description><generator>Hugo</generator><language>ko-kr</language><lastBuildDate>Thu, 10 Sep 2026 09:30:02 +0900</lastBuildDate><atom:link href="https://jaehun.me/tags/distributed-computing/index.xml" rel="self" type="application/rss+xml"/><item><title>Miles v0.1: Production-Level Post-Training</title><link>https://jaehun.me/posts/miles-v0.1-production-level-post-training/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/miles-v0.1-production-level-post-training/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2609.08368v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h2 id="miles-v01-프론티어-rl-사후학습을-위한-검증청결확장-가능한-풀스택-시스템"&gt;Miles v0.1: 프론티어 RL 사후학습을 위한 &amp;ldquo;검증·청결·확장 가능&amp;quot;한 풀스택 시스템&lt;a href="#miles-v01-%ed%94%84%eb%a1%a0%ed%8b%b0%ec%96%b4-rl-%ec%82%ac%ed%9b%84%ed%95%99%ec%8a%b5%ec%9d%84-%ec%9c%84%ed%95%9c-%ea%b2%80%ec%a6%9d%ec%b2%ad%ea%b2%b0%ed%99%95%ec%9e%a5-%ea%b0%80%eb%8a%a5%ed%95%9c-%ed%92%80%ec%8a%a4%ed%83%9d-%ec%8b%9c%ec%8a%a4%ed%85%9c" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; : Miles는 rolluout 생성과 트레이닝이 서로 다른 엔진·커널·정밀도로 돌아가며 발생하는 &lt;strong&gt;train-rollout mismatch&lt;/strong&gt;를 시스템 차원에서 해결하는 풀스택 RL 사후학습 프레임워크다. SGLang 기반 롤아웃, Megatron-LM/FSDP 트레이너, 세 가지 가중치 동기화 수송(broadcast/P2P/disk-delta)을 하나로 묶고, 토큰 정확성(TITO)과 전문가 라우팅 재현(R3)까지 보장한다. 종단 사례 연구로 &lt;strong&gt;GLM-5.2 744B-A40B&lt;/strong&gt; 를 64개 GB300 GPU에서 완전 비동기 에이전틱 RL로 학습시켜 &lt;strong&gt;중앙값 스텝 263초&lt;/strong&gt; 를 달성했다(근거: §9.2, Fig. 5).&lt;/p&gt;</description></item><item><title>Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training</title><link>https://jaehun.me/posts/online-draft-co-training-for-speculative-decoding-in-large-scale-long-context-rl-post-training/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/online-draft-co-training-for-speculative-decoding-in-large-scale-long-context-rl-post-training/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2609.07108v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h2 id="대규모장문맥-rl-사후학습에서-speculative-decoding을-실용화하는-두-가지-설계"&gt;대규모·장문맥 RL 사후학습에서 Speculative Decoding을 실용화하는 두 가지 설계&lt;a href="#%eb%8c%80%ea%b7%9c%eb%aa%a8%ec%9e%a5%eb%ac%b8%eb%a7%a5-rl-%ec%82%ac%ed%9b%84%ed%95%99%ec%8a%b5%ec%97%90%ec%84%9c-speculative-decoding%ec%9d%84-%ec%8b%a4%ec%9a%a9%ed%99%94%ed%95%98%eb%8a%94-%eb%91%90-%ea%b0%80%ec%a7%80-%ec%84%a4%ea%b3%84" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; : RL 사후학습의 벽-클록 시간을 지배하는 롤아웃 생성을 speculative decoding으로 가속하고, 여기에 &lt;strong&gt;온라인 드래프트 공동학습(co-training)&lt;/strong&gt; 을 더해 정책이 진화해도 드래프트가 뒤처지지 않게 한다. 다만 기존 시스템은 고급 드래프트가 쓰는 &lt;strong&gt;branch attention(문맥 병렬화 미지원)&lt;/strong&gt; 과 &lt;strong&gt;여러 스테이지에 흩어진 target feature(파이프라인 병렬화 문제)&lt;/strong&gt; 를 다루지 못했다. 이 논문은 (1) branch attention을 causal 주시퀀스 + rank-local branch로 분해해 병합하는 CP 기법과, (2) 파이프라인 스케줄 밖에서 target feature를 전달하는 &lt;strong&gt;TapChannel&lt;/strong&gt; 을 제안해, 8B부터 122B까지 &lt;strong&gt;1.16–1.88×&lt;/strong&gt; 의 end-to-end 가속을 달성했다(근거: §3.3, Tab. 1).&lt;/p&gt;</description></item><item><title>PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters</title><link>https://jaehun.me/posts/prima.cpp-speeding-up-70b-scale-llm-inference-on-low-resource-everyday-home-clusters/</link><pubDate>Thu, 19 Jun 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/prima.cpp-speeding-up-70b-scale-llm-inference-on-low-resource-everyday-home-clusters/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2504.08791v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="가정용-기기로-700억-매개변수-llm을-구동하다-primacpp-논문-심층-리뷰"&gt;가정용 기기로 700억 매개변수 LLM을 구동하다: prima.cpp 논문 심층 리뷰&lt;a href="#%ea%b0%80%ec%a0%95%ec%9a%a9-%ea%b8%b0%ea%b8%b0%eb%a1%9c-700%ec%96%b5-%eb%a7%a4%ea%b0%9c%eb%b3%80%ec%88%98-llm%ec%9d%84-%ea%b5%ac%eb%8f%99%ed%95%98%eb%8b%a4-primacpp-%eb%85%bc%eb%ac%b8-%ec%8b%ac%ec%b8%b5-%eb%a6%ac%eb%b7%b0" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;최근 대규모 언어 모델(LLM)의 발전은 놀랍지만, 대부분의 강력한 모델은 막대한 자원을 갖춘 클라우드 데이터센터에서만 접근 가능했습니다. GPT-4, Claude 3.5와 같은 모델을 개인 기기에서 사용하는 것은 먼 꿈처럼 여겨졌죠. 하지만 만약 여러분의 가정에 있는 노트북, 데스크톱, 스마트폰, 태블릿을 하나로 묶어 거대한 LLM을 구동할 수 있다면 어떨까요?&lt;/p&gt;</description></item><item><title>Accelerating MoE Model Inference with Expert Sharding</title><link>https://jaehun.me/posts/accelerating-moe-model-inference-with-expert-sharding/</link><pubDate>Thu, 05 Jun 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/accelerating-moe-model-inference-with-expert-sharding/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2503.08467v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;제공해주신 논문 &amp;ldquo;Accelerating MoE Model Inference with Expert Sharding&amp;rdquo; (MOESHARD)에 대한 자세한 분석은 다음과 같습니다.&lt;/p&gt;</description></item><item><title>Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures</title><link>https://jaehun.me/posts/insights-into-deepseek-v3-scaling-challenges-and-reflections-on-hardware-for-ai-architectures/</link><pubDate>Sat, 17 May 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/insights-into-deepseek-v3-scaling-challenges-and-reflections-on-hardware-for-ai-architectures/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2505.09343v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="deepseek-v3-논문-심층-분석-확장성-문제-해결-및-ai-하드웨어-아키텍처에-대한-성찰"&gt;DeepSeek-V3 논문 심층 분석: 확장성 문제 해결 및 AI 하드웨어 아키텍처에 대한 성찰&lt;a href="#deepseek-v3-%eb%85%bc%eb%ac%b8-%ec%8b%ac%ec%b8%b5-%eb%b6%84%ec%84%9d-%ed%99%95%ec%9e%a5%ec%84%b1-%eb%ac%b8%ec%a0%9c-%ed%95%b4%ea%b2%b0-%eb%b0%8f-ai-%ed%95%98%eb%93%9c%ec%9b%a8%ec%96%b4-%ec%95%84%ed%82%a4%ed%85%8d%ec%b2%98%ec%97%90-%eb%8c%80%ed%95%9c-%ec%84%b1%ec%b0%b0" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;DeepSeek-AI에서 발표한 &amp;ldquo;DeepSeek-V3: 확장성 문제 해결 및 AI 하드웨어 아키텍처에 대한 성찰&amp;rdquo; 논문은 대규모 언어 모델(LLM)의 급격한 확장에 따른 현재 하드웨어 아키텍처의 한계를 분석하고, DeepSeek-V3 모델을 통해 이러한 문제점을 효과적으로 해결하는 방안을 제시합니다. 본 논문은 하드웨어 인식 모델 공동 설계를 통해 비용 효율적인 대규모 학습 및 추론을 가능하게 하는 혁신적인 접근 방식을 상세히 설명합니다. [cite: 1, 2, 3, 4]&lt;/p&gt;</description></item><item><title>Seesaw: High-throughput LLM Inference via Model Re-sharding</title><link>https://jaehun.me/posts/seesaw-high-throughput-llm-inference-via-model-re-sharding/</link><pubDate>Mon, 12 May 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/seesaw-high-throughput-llm-inference-via-model-re-sharding/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2503.06433v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="-결론-요약"&gt;📌 결론 요약&lt;a href="#-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;논문 ‘Seesaw: High-throughput LLM Inference via Model Re-sharding’은 LLM의 두 주요 단계인 prefill과 decode에서 병렬화 전략을 동적으로 변경하는 ‘model re-sharding’ 기법을 제안하여, 평균 1.36배, 최대 1.78배의 추론 처리량 개선을 달성&lt;/strong&gt;합니다. 이 방식은 기존 vLLM처럼 고정된 병렬화 전략에 비해 throughput 최적화를 달성하며, tiered KV cache buffering과 transition-minimizing scheduling을 통해 재샤딩 비용까지 효과적으로 줄입니다.&lt;/p&gt;</description></item><item><title>Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts</title><link>https://jaehun.me/posts/comet-fine-grained-computation-communication-overlapping-for-mixture-of-experts/</link><pubDate>Mon, 14 Apr 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/comet-fine-grained-computation-communication-overlapping-for-mixture-of-experts/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.19811v3"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문 **&amp;ldquo;Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts&amp;rdquo;**는 MoE 모델의 핵심 병목인 GPU 간 통신 지연을 &lt;strong&gt;fine-grained 수준에서 컴퓨팅과 통신을 정교하게 겹치도록 설계&lt;/strong&gt;함으로써 실행 성능을 크게 개선한 ByteDance의 시스템 최적화 논문입니다.&lt;/p&gt;</description></item><item><title>MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism</title><link>https://jaehun.me/posts/megascale-infer-serving-mixture-of-experts-at-scale-with-disaggregated-expert-parallelism/</link><pubDate>Mon, 14 Apr 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/megascale-infer-serving-mixture-of-experts-at-scale-with-disaggregated-expert-parallelism/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2504.02263v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="-결론-요약-핵심-기여-및-성능"&gt;📌 결론 요약 (핵심 기여 및 성능)&lt;a href="#-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd-%ed%95%b5%ec%8b%ac-%ea%b8%b0%ec%97%ac-%eb%b0%8f-%ec%84%b1%eb%8a%a5" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;&lt;strong&gt;MegaScale-Infer&lt;/strong&gt;는 대규모 Mixture-of-Experts (MoE) 모델 서빙을 위한 효율적 시스템으로, &lt;strong&gt;Attention과 FFN 모듈을 분리(disaggregate)&lt;/strong&gt; 하여 GPU 활용률을 극대화하고 &lt;strong&gt;최대 1.9×의 GPU throughput 개선&lt;/strong&gt; 및 &lt;strong&gt;1.86× 비용 대비 성능 향상&lt;/strong&gt;을 달성합니다.&lt;/p&gt;</description></item><item><title>AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds</title><link>https://jaehun.me/posts/aiopslab-a-holistic-framework-to-evaluate-ai-agents-for-enabling-autonomous-clouds/</link><pubDate>Wed, 02 Apr 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/aiopslab-a-holistic-framework-to-evaluate-ai-agents-for-enabling-autonomous-clouds/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2501.06706"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문 『AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds』를 매우 상세히 분석한 결과, 다음과 같은 핵심 사항과 독창적 특징을 확인할 수 있었습니다.&lt;/p&gt;</description></item><item><title>NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference</title><link>https://jaehun.me/posts/neo-saving-gpu-memory-crisis-with-cpu-offloading-for-online-llm-inference/</link><pubDate>Mon, 31 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/neo-saving-gpu-memory-crisis-with-cpu-offloading-for-online-llm-inference/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2411.01142"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="-논문-제목-neo-saving-gpu-memory-crisis-with-cpu-offloading-for-online-llm-inference"&gt;📌 &lt;strong&gt;논문 제목:&lt;/strong&gt; NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference&lt;a href="#-%eb%85%bc%eb%ac%b8-%ec%a0%9c%eb%aa%a9-neo-saving-gpu-memory-crisis-with-cpu-offloading-for-online-llm-inference" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;h3 id="-저자-xuanlin-jiang-yang-zhou-shiyi-cao-ion-stoica-minlan-yu"&gt;📌 &lt;strong&gt;저자:&lt;/strong&gt; Xuanlin Jiang, Yang Zhou, Shiyi Cao, Ion Stoica, Minlan Yu&lt;a href="#-%ec%a0%80%ec%9e%90-xuanlin-jiang-yang-zhou-shiyi-cao-ion-stoica-minlan-yu" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;hr&gt;&#10;&lt;h2 id="1-결론-요약-강점--독창적인-지점"&gt;&lt;strong&gt;1. 결론 요약 (강점 &amp;amp; 독창적인 지점)&lt;/strong&gt;&lt;a href="#1-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd-%ea%b0%95%ec%a0%90--%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;NEO는 GPU 메모리의 제한으로 인해 발생하는 LLM 추론의 병목을 해결하기 위해 &lt;strong&gt;비대칭 GPU-CPU 파이프라이닝과 부하 인식 스케줄링을 적용한 새로운 시스템&lt;/strong&gt;입니다. 주요 강점은 다음과 같습니다.&lt;/p&gt;</description></item><item><title>PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training</title><link>https://jaehun.me/posts/pipefill-using-gpus-during-bubbles-in-pipeline-parallel-llm-training/</link><pubDate>Tue, 25 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/pipefill-using-gpus-during-bubbles-in-pipeline-parallel-llm-training/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2410.07192"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문『PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training』의 핵심 내용을 상세히 분석하여, 논문의 강점, 독창적인 지점, 핵심 알고리즘의 전체적인 과정 및 한계점을 요약하였습니다.&lt;/p&gt;</description></item><item><title>On Distributed Larger-Than-Memory Subset Selection With Pairwise Submodular Functions</title><link>https://jaehun.me/posts/on-distributed-larger-than-memory-subset-selection-with-pairwise-submodular-functions/</link><pubDate>Mon, 24 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/on-distributed-larger-than-memory-subset-selection-with-pairwise-submodular-functions/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2402.16442"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="-결론-요약"&gt;✅ 결론 요약&lt;a href="#-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문은 &amp;ldquo;메모리 용량을 초과하는 대규모 데이터셋에서 대표적인 subset을 효율적으로 선택&amp;quot;하는 문제를 다룬다. 기존 방법들은 중앙 서버가 전체 subset을 메모리에 올릴 수 있어야 한다는 제약이 있었지만, 이 논문은 &lt;strong&gt;중앙 서버 없이도 분산 환경에서 고품질 subset을 선택&lt;/strong&gt;할 수 있는 새로운 알고리즘 2가지를 제안한다:&lt;/p&gt;</description></item><item><title>TRAINING ULTRA LONG CONTEXT LANGUAGE MODEL WITH FULLY PIPELINED DISTRIBUTED TRANSFORMER</title><link>https://jaehun.me/posts/training-ultra-long-context-language-model-with-fully-pipelined-distributed-transformer/</link><pubDate>Mon, 24 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/training-ultra-long-context-language-model-with-fully-pipelined-distributed-transformer/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2408.16978"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="-결론-요약"&gt;✅ 결론 요약&lt;a href="#-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문은 &lt;strong&gt;초장문(long-context)&lt;/strong&gt; LLM을 &lt;strong&gt;저렴한 하드웨어(예: 4 GPU)&lt;/strong&gt; 상에서 효율적으로 훈련할 수 있게 하는 &lt;strong&gt;FPDT (Fully Pipelined Distributed Transformer)&lt;/strong&gt; 구조를 제안함.&lt;br&gt;&#10;기존 대비 &lt;strong&gt;최대 16배 더 긴 시퀀스&lt;/strong&gt;(예: 2M tokens)를 &lt;strong&gt;55% 이상의 MFU(Model FLOPs Utilization)&lt;/strong&gt; 효율로 훈련 가능함.&lt;/p&gt;</description></item><item><title>Venn: Resource Management Across Federated Learning Jobs</title><link>https://jaehun.me/posts/venn-resource-management-across-federated-learning-jobs/</link><pubDate>Tue, 18 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/venn-resource-management-across-federated-learning-jobs/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2312.08298"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="논문의-강점과-독창적인-지점"&gt;논문의 강점과 독창적인 지점&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;이 논문은 &lt;strong&gt;연합 학습(Federated Learning, FL)의 리소스 관리 문제&lt;/strong&gt;를 다루며, 특히 다수의 FL 작업이 동일한 디바이스 풀에서 실행될 때 발생하는 &lt;strong&gt;자원 경쟁(Resource Contention)&lt;/strong&gt; 을 해결하는 &lt;strong&gt;Venn&lt;/strong&gt;이라는 새로운 리소스 관리 시스템을 제안합니다.&lt;/p&gt;</description></item><item><title>AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution</title><link>https://jaehun.me/posts/ai-metropolis-scaling-large-language-model-based-multi-agent-simulation-with-out-of-order-execution/</link><pubDate>Mon, 17 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/ai-metropolis-scaling-large-language-model-based-multi-agent-simulation-with-out-of-order-execution/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2411.03519"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문『AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution』의 주요 강점과 독창적인 지점, 핵심 알고리즘 및 한계점을 압축하여 설명하면 다음과 같습니다.&lt;/p&gt;</description></item><item><title>Balancing Pipeline Parallelism with Vocabulary Parallelism</title><link>https://jaehun.me/posts/balancing-pipeline-parallelism-with-vocabulary-parallelism/</link><pubDate>Mon, 17 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/balancing-pipeline-parallelism-with-vocabulary-parallelism/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2411.05288"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문의 핵심 내용을 요약하면 다음과 같습니다.&lt;/p&gt;</description></item><item><title>DIFFSERVE: EFFICIENTLY SERVING TEXT-TO-IMAGE DIFFUSION MODELS WITH QUERY-AWARE MODEL SCALING</title><link>https://jaehun.me/posts/diffserve-efficiently-serving-text-to-image-diffusion-models-with-query-aware-model-scaling/</link><pubDate>Mon, 17 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/diffserve-efficiently-serving-text-to-image-diffusion-models-with-query-aware-model-scaling/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2411.15381"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-핵심-요약-및-기여"&gt;&lt;strong&gt;논문의 핵심 요약 및 기여&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%ec%9a%94%ec%95%bd-%eb%b0%8f-%ea%b8%b0%ec%97%ac" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;DIFFSERVE&lt;/strong&gt;는 &lt;strong&gt;query-aware model scaling&lt;/strong&gt; 개념을 도입하여 &lt;strong&gt;Text-to-Image Diffusion Model&lt;/strong&gt;의 효율적인 서빙을 가능하게 하는 시스템이다. 기존 서빙 시스템이 모든 요청에 대해 동일한 크기의 모델을 사용하는 반면, DIFFSERVE는 입력 쿼리의 난이도에 따라 &lt;strong&gt;경량(lightweight) 모델과 고성능(heavyweight) 모델을 선택적으로 사용&lt;/strong&gt;하는 &lt;strong&gt;모델 캐스케이드(model cascade)&lt;/strong&gt; 기법을 적용한다. 이를 통해 &lt;strong&gt;최대 24% 품질 향상&lt;/strong&gt;, &lt;strong&gt;19-70% SLO(서비스 레벨 목표) 위반 감소&lt;/strong&gt;를 달성한다.&lt;/p&gt;</description></item><item><title>Marconi: Prefix Caching for the Era of Hybrid LLMs</title><link>https://jaehun.me/posts/marconi-prefix-caching-for-the-era-of-hybrid-llms/</link><pubDate>Wed, 12 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/marconi-prefix-caching-for-the-era-of-hybrid-llms/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2411.19379"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-핵심-요약-및-평가"&gt;논문의 핵심 요약 및 평가&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%ec%9a%94%ec%95%bd-%eb%b0%8f-%ed%8f%89%ea%b0%80" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;논문 제목:&lt;/strong&gt;&lt;br&gt;&#10;&lt;strong&gt;Marconi: Prefix Caching for the Era of Hybrid LLMs&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>LAVA: LIFETIME-AWARE VM ALLOCATION WITH LEARNED DISTRIBUTIONS AND ADAPTATION TO MISPREDICTIONS</title><link>https://jaehun.me/posts/lava-lifetime-aware-vm-allocation-with-learned-distributions-and-adaptation-to-mispredictions/</link><pubDate>Tue, 11 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/lava-lifetime-aware-vm-allocation-with-learned-distributions-and-adaptation-to-mispredictions/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2412.09840"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="논문의-강점과-독창성"&gt;&lt;strong&gt;논문의 강점과 독창성&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%84%b1" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;h3 id="강점"&gt;&lt;strong&gt;강점&lt;/strong&gt;&lt;a href="#%ea%b0%95%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;기존 VM 스케줄링 방식보다 높은 효율성&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>A PRACTICAL CROSS-LAYER APPROACH FOR ML-DRIVEN STORAGE PLACEMENT IN WAREHOUSE-SCALE COMPUTERS</title><link>https://jaehun.me/posts/a-practical-cross-layer-approach-for-ml-driven-storage-placement-in-warehouse-scale-computers/</link><pubDate>Mon, 10 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/a-practical-cross-layer-approach-for-ml-driven-storage-placement-in-warehouse-scale-computers/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2501.05651"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-및-독창적인-지점"&gt;&lt;strong&gt;논문의 강점 및 독창적인 지점&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문은 &lt;strong&gt;대규모 데이터 센터에서 기계 학습(ML)을 활용한 저장소 배치(Storage Placement) 문제&lt;/strong&gt;를 다루며, 기존 접근법의 한계를 극복하기 위해 &lt;strong&gt;크로스-레이어(Cross-Layer) 접근 방식&lt;/strong&gt;을 제안했다. 주요 강점과 독창적인 점은 다음과 같다.&lt;/p&gt;</description></item><item><title>Scaling Deep Learning Training with MPMD Pipeline Parallelism</title><link>https://jaehun.me/posts/scaling-deep-learning-training-with-mpmd-pipeline-parallelism/</link><pubDate>Mon, 10 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/scaling-deep-learning-training-with-mpmd-pipeline-parallelism/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2412.14374"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;본 논문은 JaxPP라는 시스템을 제안하여, 기존의 Single-Program Multiple-Data (SPMD) 방식의 한계를 극복하고 Multiple-Program Multiple-Data (MPMD) 파이프라인 병렬화를 통해 대규모 딥러닝 모델 학습의 확장성과 성능을 향상한 연구이다. 특히, JaxPP는 사용자가 pipeline 스케줄링을 유연하게 정의할 수 있도록 지원하며, 자동화된 작업 분배와 통신 패턴 추론을 통해 하드웨어 자원을 효율적으로 사용하여 기존 SPMD 대비 최대 1.11배 향상된 성능을 보였다.&lt;/p&gt;</description></item><item><title>LSERVE: EFFICIENT LONG-SEQUENCE LLM SERVING WITH UNIFIED SPARSE ATTENTION</title><link>https://jaehun.me/posts/lserve-efficient-long-sequence-llm-serving-with-unified-sparse-attention/</link><pubDate>Thu, 06 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/lserve-efficient-long-sequence-llm-serving-with-unified-sparse-attention/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.14866"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문의 핵심 내용을 먼저 간략히 요약한 후, 강점과 독창적인 지점을 자세히 설명하고, 핵심 알고리즘의 동작 원리를 예시와 함께 제시한 뒤, 논문의 한계점을 마지막으로 정리하겠습니다.&lt;/p&gt;</description></item><item><title>ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments</title><link>https://jaehun.me/posts/thunderserve-high-performance-and-cost-efficient-llm-serving-in-cloud-environments/</link><pubDate>Thu, 06 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/thunderserve-high-performance-and-cost-efficient-llm-serving-in-cloud-environments/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.09334"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문 『ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments』를 상세히 분석하여 다음과 같은 내용을 압축적으로 정리하였다.&lt;/p&gt;</description></item><item><title>HEXGEN-2: DISAGGREGATED GENERATIVE INFERENCE OF LLMS IN HETEROGENEOUS ENVIRONMENT</title><link>https://jaehun.me/posts/hexgen-2-disaggregated-generative-inference-of-llms-in-heterogeneous-environment/</link><pubDate>Tue, 25 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/hexgen-2-disaggregated-generative-inference-of-llms-in-heterogeneous-environment/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.07903v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-핵심-요약-및-기여점"&gt;&lt;strong&gt;논문의 핵심 요약 및 기여점&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%ec%9a%94%ec%95%bd-%eb%b0%8f-%ea%b8%b0%ec%97%ac%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문은 &lt;strong&gt;HEXGEN-2&lt;/strong&gt;라는 새로운 분산 LLM(대형 언어 모델) 추론 시스템을 제안합니다. 기존 동질적인 고성능 GPU 클러스터를 이용하는 방식과 달리, &lt;strong&gt;이기종 GPU 환경&lt;/strong&gt;에서 &lt;strong&gt;Prefill(입력 처리)과 Decoding(출력 생성) 단계를 분리(disaggregated inference)&lt;/strong&gt; 하여 비용 효율성을 극대화하는 것이 핵심 아이디어입니다.&lt;/p&gt;</description></item><item><title>A Hardware Evaluation Framework for Large Language Model Inference</title><link>https://jaehun.me/posts/a-hardware-evaluation-framework-for-large-language-model-inference/</link><pubDate>Tue, 21 Jan 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/a-hardware-evaluation-framework-for-large-language-model-inference/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2312.03134"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-및-독창적인-지점"&gt;논문의 강점 및 독창적인 지점&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;빠르고 정확한 하드웨어 평가 프레임워크 (LLMCompass) 제공&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Fast State Restoration in LLM Serving with HCache</title><link>https://jaehun.me/posts/fast-state-restoration-in-llm-serving-with-hcache/</link><pubDate>Tue, 21 Jan 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/fast-state-restoration-in-llm-serving-with-hcache/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2410.05004v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="논문의-강점-및-독창성"&gt;논문의 강점 및 독창성&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%84%b1" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;h3 id="강점"&gt;강점&lt;a href="#%ea%b0%95%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;효율적인 상태 복원 기술 (HCache) 제안&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication</title><link>https://jaehun.me/posts/tokenring-an-efficient-parallelism-framework-for-infinite-context-llms-via-bidirectional-communication/</link><pubDate>Mon, 20 Jan 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/tokenring-an-efficient-parallelism-framework-for-infinite-context-llms-via-bidirectional-communication/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2412.20501v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-및-독창성"&gt;논문의 강점 및 독창성&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%84%b1" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;강점:&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Distributed Inference and Fine-tuning of Large Language Models Over The Internet</title><link>https://jaehun.me/posts/distributed-inference-and-fine-tuning-of-large-language-models-over-the-internet/</link><pubDate>Thu, 02 Jan 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/distributed-inference-and-fine-tuning-of-large-language-models-over-the-internet/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2312.08361"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점과-독창적인-지점"&gt;논문의 강점과 독창적인 지점&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;독창적인 분산 시스템 설계&lt;/strong&gt;:&lt;/p&gt;</description></item><item><title>EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism</title><link>https://jaehun.me/posts/ee-llm-large-scale-training-and-inference-of-early-exit-large-language-models-with-3d-parallelism/</link><pubDate>Thu, 02 Jan 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/ee-llm-large-scale-training-and-inference-of-early-exit-large-language-models-with-3d-parallelism/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2312.04916"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-주요-강점과-독창성"&gt;논문의 주요 강점과 독창성&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ec%a3%bc%ec%9a%94-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%84%b1" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;효율적인 대규모 LLM 훈련 및 추론&lt;/strong&gt;:&lt;/p&gt;</description></item><item><title>Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving</title><link>https://jaehun.me/posts/mooncake-a-kvcache-centric-disaggregated-architecture-for-llm-serving/</link><pubDate>Tue, 24 Dec 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/mooncake-a-kvcache-centric-disaggregated-architecture-for-llm-serving/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2407.00079"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;&lt;strong&gt;Mooncake: A KVCache-Centric Disaggregated Architecture for LLM Serving&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>