<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hardware Architecture on Jaehun's Blog</title><link>https://jaehun.me/tags/hardware-architecture/</link><description>Recent content in Hardware Architecture on Jaehun's Blog</description><generator>Hugo</generator><language>ko-kr</language><lastBuildDate>Fri, 18 Sep 2026 10:25:57 +0900</lastBuildDate><atom:link href="https://jaehun.me/tags/hardware-architecture/index.xml" rel="self" type="application/rss+xml"/><item><title>The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction</title><link>https://jaehun.me/posts/paper-2609-18063/</link><pubDate>Fri, 18 Sep 2026 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/paper-2609-18063/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2609.18063"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h2 id="ssd에-산-moe를-깨우다-24-gb-데스크톱에서-35b-moe를-20-toks로-서빙하는-edge0"&gt;SSD에 산 MoE를 깨우다: 24 GB 데스크톱에서 35B MoE를 20 tok/s로 서빙하는 Edge0&lt;a href="#ssd%ec%97%90-%ec%82%b0-moe%eb%a5%bc-%ea%b9%a8%ec%9a%b0%eb%8b%a4-24-gb-%eb%8d%b0%ec%8a%a4%ed%81%ac%ed%86%b1%ec%97%90%ec%84%9c-35b-moe%eb%a5%bc-20-toks%eb%a1%9c-%ec%84%9c%eb%b9%99%ed%95%98%eb%8a%94-edge0" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;h2 id="한-줄-요약-tldr"&gt;한 줄 요약 (TL;DR)&lt;a href="#%ed%95%9c-%ec%a4%84-%ec%9a%94%ec%95%bd-tldr" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;35B급 MoE는 4-bit로도 19.5 GB 를 차지해 24 GB 머신에 상주하지 못한다 (근거: §1). Edge0는 엑스퍼트 가중치를 SSD에 두고 mmap 스트리밍으로 읽되, 다음 레이어 라우팅을 한 토큰 앞서 예측하는 프리라우터(prerouter)로 디스크 지연을 컴퓨트 뒤에 숨기고, 그 예측 자체를 라우팅으로 소비해 드롭을 없애며, int4 + 라우팅 교체 손실을 병합하지 않은(unmerged) 복구 LoRA로 되갚는다 (근거: §3.2, §3.3). 결과는 Mac mini M4 Pro 24 GB 에서 35B 티어 20.4 tok/s , 피크 활성 메모리 2.9 GiB , fp16 교사 대비 평균 3.9 포인트 갭이다 (근거: Tab.1, Tab.2).&lt;/p&gt;</description></item><item><title>Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures</title><link>https://jaehun.me/posts/insights-into-deepseek-v3-scaling-challenges-and-reflections-on-hardware-for-ai-architectures/</link><pubDate>Sat, 17 May 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/insights-into-deepseek-v3-scaling-challenges-and-reflections-on-hardware-for-ai-architectures/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2505.09343v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="deepseek-v3-논문-심층-분석-확장성-문제-해결-및-ai-하드웨어-아키텍처에-대한-성찰"&gt;DeepSeek-V3 논문 심층 분석: 확장성 문제 해결 및 AI 하드웨어 아키텍처에 대한 성찰&lt;a href="#deepseek-v3-%eb%85%bc%eb%ac%b8-%ec%8b%ac%ec%b8%b5-%eb%b6%84%ec%84%9d-%ed%99%95%ec%9e%a5%ec%84%b1-%eb%ac%b8%ec%a0%9c-%ed%95%b4%ea%b2%b0-%eb%b0%8f-ai-%ed%95%98%eb%93%9c%ec%9b%a8%ec%96%b4-%ec%95%84%ed%82%a4%ed%85%8d%ec%b2%98%ec%97%90-%eb%8c%80%ed%95%9c-%ec%84%b1%ec%b0%b0" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;DeepSeek-AI에서 발표한 &amp;ldquo;DeepSeek-V3: 확장성 문제 해결 및 AI 하드웨어 아키텍처에 대한 성찰&amp;rdquo; 논문은 대규모 언어 모델(LLM)의 급격한 확장에 따른 현재 하드웨어 아키텍처의 한계를 분석하고, DeepSeek-V3 모델을 통해 이러한 문제점을 효과적으로 해결하는 방안을 제시합니다. 본 논문은 하드웨어 인식 모델 공동 설계를 통해 비용 효율적인 대규모 학습 및 추론을 가능하게 하는 혁신적인 접근 방식을 상세히 설명합니다. [cite: 1, 2, 3, 4]&lt;/p&gt;</description></item><item><title>Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching</title><link>https://jaehun.me/posts/duplex-a-device-for-large-language-models-with-mixture-of-experts-grouped-query-attention-and-continuous-batching/</link><pubDate>Sun, 13 Apr 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/duplex-a-device-for-large-language-models-with-mixture-of-experts-grouped-query-attention-and-continuous-batching/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2409.01141v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="-결론-요약"&gt;📌 결론 요약&lt;a href="#-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;논문 *&amp;ldquo;Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching&amp;rdquo;*은 저연산량(Op/B)이 지배적인 MoE 및 GQA 기반 LLM 추론을 위한 하드웨어 아키텍처 &lt;strong&gt;Duplex&lt;/strong&gt;를 제안하며, GPU 단독 대비 &lt;strong&gt;최대 2.67×의 추론 속도&lt;/strong&gt;와 &lt;strong&gt;42.03%의 에너지 절감&lt;/strong&gt; 효과를 보여줍니다. 핵심은 **xPU (GPU 수준 고성능 연산기)**와 **Logic-PIM (로직 다이에 탑재된 저 Op/B 특화 연산기)**를 &lt;strong&gt;동시에 활용&lt;/strong&gt;하여 MoE와 Attention Layer를 공동 처리(co-processing)하는 방식입니다.&lt;/p&gt;</description></item><item><title>LeanAttention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers</title><link>https://jaehun.me/posts/leanattention-hardware-aware-scalable-attention-mechanism-for-the-decode-phase-of-transformers/</link><pubDate>Mon, 07 Apr 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/leanattention-hardware-aware-scalable-attention-mechanism-for-the-decode-phase-of-transformers/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2405.10480v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문 『LeanAttention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers』에 대한 상세한 분석을 다음과 같이 제시합니다.&lt;/p&gt;</description></item><item><title>TurboAttention: Efficient Attention Approximation for High Throughputs LLMs</title><link>https://jaehun.me/posts/turboattention-efficient-attention-approximation-for-high-throughputs-llms/</link><pubDate>Tue, 11 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/turboattention-efficient-attention-approximation-for-high-throughputs-llms/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2412.08585"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문 『TurboAttention: Efficient Attention Approximation for High Throughputs LLMs』는 기존의 Attention 연산의 속도와 메모리 효율성을 동시에 개선한 통합적인 접근 방법을 제안하고 있으며, 이는 두 가지 핵심 알고리즘인 FlashQ와 Sparsity-based Softmax Approximation (SAS)을 통해 구현되었습니다.&lt;/p&gt;</description></item><item><title>A Hardware Evaluation Framework for Large Language Model Inference</title><link>https://jaehun.me/posts/a-hardware-evaluation-framework-for-large-language-model-inference/</link><pubDate>Tue, 21 Jan 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/a-hardware-evaluation-framework-for-large-language-model-inference/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2312.03134"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-및-독창적인-지점"&gt;논문의 강점 및 독창적인 지점&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;빠르고 정확한 하드웨어 평가 프레임워크 (LLMCompass) 제공&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving</title><link>https://jaehun.me/posts/mooncake-a-kvcache-centric-disaggregated-architecture-for-llm-serving/</link><pubDate>Tue, 24 Dec 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/mooncake-a-kvcache-centric-disaggregated-architecture-for-llm-serving/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2407.00079"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;&lt;strong&gt;Mooncake: A KVCache-Centric Disaggregated Architecture for LLM Serving&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache</title><link>https://jaehun.me/posts/infinite-llm-efficient-llm-service-for-long-context-with-distattention-and-distributed-kvcache/</link><pubDate>Wed, 18 Dec 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/infinite-llm-efficient-llm-service-for-long-context-with-distattention-and-distributed-kvcache/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2401.02669"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문-요약"&gt;논문 요약&lt;a href="#%eb%85%bc%eb%ac%b8-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;Infinite-LLM&lt;/strong&gt;은 초대형 언어 모델(LLM)의 동적 컨텍스트 길이 문제를 해결하기 위해 제안된 효율적인 서비스 시스템입니다. 논문은 특히 &lt;strong&gt;DistAttention&lt;/strong&gt; 메커니즘과 &lt;strong&gt;Distributed KVCache&lt;/strong&gt;를 도입해 LLM 요청 서비스의 확장성을 획기적으로 개선합니다. 이를 통해 긴 컨텍스트 길이를 효과적으로 처리하며, 시스템 전체의 계산 및 메모리 자원 활용도를 높입니다.&lt;/p&gt;</description></item><item><title>FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs</title><link>https://jaehun.me/posts/flightllm-efficient-large-language-model-inference-with-a-complete-mapping-flow-on-fpgas/</link><pubDate>Tue, 17 Dec 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/flightllm-efficient-large-language-model-inference-with-a-complete-mapping-flow-on-fpgas/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2401.03868"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문-요약-및-분석"&gt;논문 요약 및 분석&lt;a href="#%eb%85%bc%eb%ac%b8-%ec%9a%94%ec%95%bd-%eb%b0%8f-%eb%b6%84%ec%84%9d" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;제목:&lt;/strong&gt; &lt;em&gt;FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs&lt;/em&gt;&lt;br&gt;&#10;&lt;strong&gt;주제:&lt;/strong&gt; FPGA 기반 대형 언어 모델(LLM) 추론을 위한 효율적 아키텍처&lt;/p&gt;</description></item><item><title>FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design</title><link>https://jaehun.me/posts/fp6-llm-efficiently-serving-large-language-models-through-fp6-centric-algorithm-system-co-design/</link><pubDate>Sun, 15 Dec 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/fp6-llm-efficiently-serving-large-language-models-through-fp6-centric-algorithm-system-co-design/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2401.14112"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점과-독창적인-지점"&gt;논문의 강점과 독창적인 지점&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;FP6 기반의 효율적인 양자화 지원&lt;/strong&gt;:&lt;/p&gt;</description></item><item><title>Benchmarking and Dissecting the Nvidia Hopper GPU Architecture</title><link>https://jaehun.me/posts/benchmarking-and-dissecting-the-nvidia-hopper-gpu-architecture/</link><pubDate>Sat, 14 Dec 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/benchmarking-and-dissecting-the-nvidia-hopper-gpu-architecture/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2402.13499v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-주요-내용-요약"&gt;논문의 주요 내용 요약&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ec%a3%bc%ec%9a%94-%eb%82%b4%ec%9a%a9-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문은 Nvidia Hopper GPU 아키텍처에 대한 심층적인 벤치마킹과 분석을 수행했습니다. 특히 Hopper 아키텍처의 다음 세 가지 주요 특징을 다루었습니다:&lt;/p&gt;</description></item><item><title>COMET Towards Partical W4A4KV4 LLMs Serving</title><link>https://jaehun.me/posts/comet-towards-partical-w4a4kv4-llms-serving/</link><pubDate>Mon, 11 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/comet-towards-partical-w4a4kv4-llms-serving/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2410.12168"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2410.12168&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;아래글은 Scholar GPT와 대화를 나눈 내용입니다.&lt;/p&gt;&#10;&lt;h1 id="q--논문의-핵심-내용과-강점-알고리즘-설명-그리고-한계점"&gt;Q : 논문의 핵심 내용과 강점, 알고리즘 설명, 그리고 한계점&lt;a href="#q--%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%ea%b0%95%ec%a0%90-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문 **&amp;ldquo;COMET: Towards Practical W4A4KV4 LLMs Serving&amp;rdquo;**는 대형 언어 모델(LLMs)을 효과적으로 배포하기 위해 양자화(quantization) 기법을 활용하여 GPU 성능을 최적화하는 방법을 제안합니다. 특히 INT4 텐서 코어를 활용해 모델의 메모리 사용량과 계산 효율성을 극대화하는 것이 핵심입니다. 이 논문에서의 주요 내용과 기여점을 설명드리겠습니다.&lt;/p&gt;</description></item><item><title>DynamoLLM Designing LLM Inference Clusters for Performance and Energy Efficiency</title><link>https://jaehun.me/posts/dynamollm-designing-llm-inference-clusters-for-performance-and-energy-efficiency/</link><pubDate>Sun, 10 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/dynamollm-designing-llm-inference-clusters-for-performance-and-energy-efficiency/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2408.00741"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2408.00741&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;아래글은 Scholar GPT와 대화를 나눈 내용입니다.&lt;/p&gt;&#10;&lt;h1 id="q--논문의-핵심-내용과-강점-알고리즘-설명-그리고-한계점"&gt;Q : 논문의 핵심 내용과 강점, 알고리즘 설명, 그리고 한계점&lt;a href="#q--%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%ea%b0%95%ec%a0%90-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문-분석"&gt;&lt;strong&gt;논문 분석: &amp;ldquo;DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency&amp;rdquo;&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8-%eb%b6%84%ec%84%9d" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;hr&gt;&#10;&lt;h3 id="1-논문의-강점-및-독창적인-지점"&gt;&lt;strong&gt;1. 논문의 강점 및 독창적인 지점&lt;/strong&gt;&lt;a href="#1-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;h4 id="강점"&gt;&lt;strong&gt;강점:&lt;/strong&gt;&lt;a href="#%ea%b0%95%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;에너지 효율적인 LLM 인프라&lt;/strong&gt;: 이 논문은 대규모 LLM 추론 클러스터에서 &lt;strong&gt;에너지 소비를 최적화&lt;/strong&gt;하고 &lt;strong&gt;운영 비용을 줄이기 위해 DynamoLLM이라는 시스템&lt;/strong&gt;을 제안합니다. 이 시스템은 LLM 추론의 다양한 특성을 활용하여 &lt;strong&gt;에너지 효율성을 53% 개선&lt;/strong&gt;하고, &lt;strong&gt;운영 비용을 61% 절감&lt;/strong&gt;하며 &lt;strong&gt;탄소 배출량을 38% 줄였습니다&lt;/strong&gt;.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;동적 재구성 능력&lt;/strong&gt;: DynamoLLM은 &lt;strong&gt;워크로드 변화에 따라 동적으로 재구성&lt;/strong&gt;할 수 있는 시스템으로, &lt;strong&gt;서버 인스턴스 수, 모델 병렬화, GPU 주파수 조정&lt;/strong&gt; 등을 자동으로 조정합니다. 이를 통해 **성능 목표(SLO)**를 충족하면서도 에너지 소비를 최소화합니다.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;다양한 LLM 요청에 대응&lt;/strong&gt;: 이 시스템은 &lt;strong&gt;입력 및 출력 토큰 길이, 모델의 특성, SLO 요구 사항&lt;/strong&gt;에 따라 서로 다른 요청을 &lt;strong&gt;다양한 서버 풀로 분류하고 최적화된 설정으로 처리&lt;/strong&gt;합니다. 이를 통해 요청 유형에 따라 &lt;strong&gt;다양한 에너지-성능 프로파일&lt;/strong&gt;을 적용하여 효율을 높입니다.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h4 id="독창적인-지점"&gt;&lt;strong&gt;독창적인 지점:&lt;/strong&gt;&lt;a href="#%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;계층적 제어 아키텍처&lt;/strong&gt;: DynamoLLM은 &lt;strong&gt;계층적 제어 시스템&lt;/strong&gt;을 도입하여 클러스터, 풀, 인스턴스 레벨에서 각각 최적화를 수행합니다. 각 레벨에서 독립적으로 조정을 수행하며, 상위 레벨에서 전달된 제약 조건을 준수하면서 에너지 효율성을 극대화합니다.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;빠른 재구성 및 스케일링&lt;/strong&gt;: 서버 인스턴스, 모델 병렬화, GPU 주파수를 &lt;strong&gt;빠르게 조정&lt;/strong&gt;하여 &lt;strong&gt;재구성 오버헤드를 최소화&lt;/strong&gt;합니다. 이를 통해 변동하는 워크로드에도 빠르게 적응할 수 있습니다.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;hr&gt;&#10;&lt;h3 id="2-핵심-알고리즘-설명-및-예시"&gt;&lt;strong&gt;2. 핵심 알고리즘 설명 및 예시&lt;/strong&gt;&lt;a href="#2-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%eb%b0%8f-%ec%98%88%ec%8b%9c" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;h4 id="입력-예시"&gt;&lt;strong&gt;입력 예시&lt;/strong&gt;&lt;a href="#%ec%9e%85%eb%a0%a5-%ec%98%88%ec%8b%9c" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;ul&gt;&#10;&lt;li&gt;사용자가 다음과 같은 요청을 보낸다고 가정합니다:&#10;&lt;div&#10; class="code-block-container border-border bg-card my-6 overflow-hidden rounded-xl border shadow-sm transition-all duration-200 ease-out hover:-translate-y-0.5 hover:shadow-md"&#10; data-code-block&#10; data-code-id="code-0"&#10; data-collapsible="true"&#10; data-default-state="expanded"&#10; data-collapsed="false"&#10; data-auto-collapse-lines="30"&#10; data-auto-collapse-height="400"&#10; data-collapsed-height="120"&#10;&gt;&#10; &#10; &lt;div&#10; class="code-block-header bg-muted/30 border-border flex items-center justify-between border-b px-4 py-3"&#10; &gt;&#10; &#10; &lt;div class="flex items-center gap-2"&gt;&#10; &lt;div class="text-muted-foreground shrink-0"&gt;&#10; &lt;svg class="h-4 w-4" fill="none" stroke="currentColor" viewBox="0 0 24 24" xmlns="http://www.w3.org/2000/svg"&gt;&#10; &lt;path stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M10 20l4-16m4 4l4 4-4 4M6 16l-4-4 4-4" /&gt;&#10;&lt;/svg&gt;&#10; &lt;/div&gt;&#10; &lt;span class="text-muted-foreground text-sm font-medium"&gt;&#10; PLAINTEXT&#10; &lt;/span&gt;&#10; &lt;/div&gt;&#10;&#10; &#10; &lt;div class="flex items-center gap-2"&gt;&#10; &lt;button&#10; class="collapse-code-btn text-muted-foreground hover:text-primary hover:bg-primary/10 focus:ring-primary/20 flex items-center gap-1.5 rounded-md px-2 py-1 text-xs font-medium transition-all duration-200 ease-out focus:ring-2 focus:outline-none"&#10; type="button"&#10; data-code-action="toggle-collapse"&#10; data-label-expand="확장"&#10; data-label-collapse="접기"&#10; title="접기"&#10; aria-label="접기"&#10; aria-controls="code-0"&#10; aria-expanded="true"&#10; &gt;&#10; &lt;span class="collapse-chevron transition-transform duration-200 ease-out"&gt;&#10; &lt;svg class="h-3 w-3" fill="none" stroke="currentColor" viewBox="0 0 24 24" xmlns="http://www.w3.org/2000/svg"&gt;&#10; &lt;path stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M19 9l-7 7-7-7" /&gt;&#10;&lt;/svg&gt;&#10; &lt;/span&gt;&#10; &lt;span class="collapse-text hidden sm:inline"&gt;접기&lt;/span&gt;&#10; &lt;/button&gt;&#10; &lt;button&#10; class="copy-code-btn text-muted-foreground hover:text-primary hover:bg-primary/10 focus:ring-primary/20 flex items-center gap-1.5 rounded-md px-2 py-1 text-xs font-medium transition-all duration-200 ease-out focus:ring-2 focus:outline-none"&#10; type="button"&#10; data-code-action="copy"&#10; data-label-copy="복사"&#10; data-label-copied="복사됨"&#10; title="복사"&#10; aria-label="복사"&#10; &gt;&#10; &lt;span class="copy-icon"&gt;&#10; &lt;svg class="h-3 w-3" fill="none" stroke="currentColor" viewBox="0 0 24 24" xmlns="http://www.w3.org/2000/svg"&gt;&#10; &lt;path stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M8 16H6a2 2 0 01-2-2V6a2 2 0 012-2h8a2 2 0 012 2v2m-6 12h8a2 2 0 002-2v-8a2 2 0 00-2-2h-8a2 2 0 00-2 2v8a2 2 0 002 2z" /&gt;&#10;&lt;/svg&gt;&#10; &lt;/span&gt;&#10; &lt;span class="copy-text hidden sm:inline"&gt;복사&lt;/span&gt;&#10; &lt;/button&gt;&#10; &lt;/div&gt;&#10; &lt;/div&gt;&#10;&#10; &#10; &lt;div class="code-block-content relative" id="code-0"&gt;&#10; &lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&amp;#34;Llama2-70B 모델을 사용하여 100,000개의 입력 토큰을 분석하고 요약을 생성해 주세요.&amp;#34;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&#10; &lt;div hidden data-code-source&gt;&amp;#34;Llama2-70B 모델을 사용하여 100,000개의 입력 토큰을 분석하고 요약을 생성해 주세요.&amp;#34;&lt;/div&gt;&#10; &#10; &lt;div&#10; class="collapse-overlay to-card/90 pointer-events-none absolute inset-0 bg-linear-to-b from-transparent via-transparent opacity-0 transition-opacity duration-300"&#10; hidden&#10; &gt;&#10; &lt;button&#10; class="collapse-overlay-btn text-muted-foreground bg-card/80 border-border/50 hover:bg-primary/10 hover:text-primary hover:border-primary/30 absolute bottom-4 left-1/2 flex -translate-x-1/2 items-center justify-center rounded-full border p-2 backdrop-blur-sm transition-all duration-200"&#10; type="button"&#10; data-code-action="expand"&#10; aria-label="확장"&#10; title="확장"&#10; &gt;&#10; &lt;svg class="h-4 w-4" fill="none" stroke="currentColor" viewBox="0 0 24 24" xmlns="http://www.w3.org/2000/svg"&gt;&#10; &lt;path stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M19 9l-7 7-7-7" /&gt;&#10;&lt;/svg&gt;&#10; &lt;/button&gt;&#10; &lt;/div&gt;&#10; &lt;/div&gt;&#10;&lt;/div&gt;&#10;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h4 id="dynamollm의-추론-과정-inference-process"&gt;&lt;strong&gt;DynamoLLM의 추론 과정 (Inference Process)&lt;/strong&gt;&lt;a href="#dynamollm%ec%9d%98-%ec%b6%94%eb%a1%a0-%ea%b3%bc%ec%a0%95-inference-process" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;p&gt;&lt;strong&gt;Step 1: 입력 분석 및 워크로드 분류&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Efficient Generative LLM Inference Using Phase Splitting</title><link>https://jaehun.me/posts/efficient-generative-llm-inference-using-phase-splitting/</link><pubDate>Tue, 05 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/efficient-generative-llm-inference-using-phase-splitting/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2311.18677v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2311.18677v2&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;아래글은 Scholar GPT와 대화를 나눈 내용입니다.&lt;/p&gt;&#10;&lt;h1 id="q--논문의-핵심-내용과-강점-알고리즘-설명-그리고-한계점"&gt;Q : 논문의 핵심 내용과 강점, 알고리즘 설명, 그리고 한계점&lt;a href="#q--%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%ea%b0%95%ec%a0%90-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h1 id="answer"&gt;Answer&lt;a href="#answer" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문 **&amp;ldquo;Splitwise: Efficient Generative LLM Inference Using Phase Splitting&amp;rdquo;**는 대형 언어 모델(LLM)의 추론을 효율적으로 수행하기 위해서 &lt;strong&gt;추론 단계를 분리하는 Splitwise&lt;/strong&gt;라는 기법을 제안합니다. 이 논문은 주로 대형 모델의 추론 요청 시 발생하는 비효율성을 개선하고자 합니다. Splitwise는 추론 과정의 &lt;strong&gt;프롬프트 계산 단계&lt;/strong&gt;와 &lt;strong&gt;토큰 생성 단계&lt;/strong&gt;를 분리하여 각 단계에 맞는 최적의 하드웨어 자원을 할당함으로써 성능과 비용을 동시에 개선하는 시스템입니다. 이 논문의 주요 강점, 독창적인 지점, 핵심 알고리즘의 예시, 그리고 한계점은 다음과 같습니다.&lt;/p&gt;</description></item><item><title>Keyformer KV Cache Reduction through Key Tokens Selection for Efficient Generative Inference</title><link>https://jaehun.me/posts/keyformer-kv-cache-reduction-through-key-tokens-selection-for-efficient-generative-inference/</link><pubDate>Thu, 31 Oct 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/keyformer-kv-cache-reduction-through-key-tokens-selection-for-efficient-generative-inference/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2403.09054"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2403.09054&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;개인의견 : 예전부터 KV cache중에 핵심이 되는 내용만 잘 선택하면 오버헤드를 줄일수 있지 않을까라고 생각했는데 핵심 토큰만 가지고 이러한 방법을 구현한 논문입니다.&#10;이러한 논문을 볼때 마다 항상 과연 정확도에 얼마나 영향을 줄까라는 의심이 생길수 밖에 없지만 이렇게 알고리즘쪽에서 경량화를 많이 해주면 좋겠네요&lt;/p&gt;</description></item></channel></rss>