<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Natural Language Processing on Jaehun's Blog</title><link>https://jaehun.me/tags/natural-language-processing/</link><description>Recent content in Natural Language Processing on Jaehun's Blog</description><generator>Hugo</generator><language>ko-kr</language><lastBuildDate>Tue, 08 Sep 2026 13:42:29 +0000</lastBuildDate><atom:link href="https://jaehun.me/tags/natural-language-processing/index.xml" rel="self" type="application/rss+xml"/><item><title>[논문리뷰]: NVIDIA Nemotron 3: Efficient and Open Intelligence</title><link>https://jaehun.me/posts/%EB%85%BC%EB%AC%B8%EB%A6%AC%EB%B7%B0-nvidia-nemotron-3-efficient-and-open-intelligence/</link><pubDate>Tue, 16 Dec 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/%EB%85%BC%EB%AC%B8%EB%A6%AC%EB%B7%B0-nvidia-nemotron-3-efficient-and-open-intelligence/</guid><description>&lt;p&gt;&lt;a&#10; href="https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-White-Paper.pdf"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h2 id="nvidia-nemotron-3-하이브리드-mambatransformer-moe로-정확도처리량-프런티어를-당기다"&gt;NVIDIA Nemotron 3: 하이브리드 Mamba–Transformer MoE로 “정확도/처리량” 프런티어를 당기다&lt;a href="#nvidia-nemotron-3-%ed%95%98%ec%9d%b4%eb%b8%8c%eb%a6%ac%eb%93%9c-mambatransformer-moe%eb%a1%9c-%ec%a0%95%ed%99%95%eb%8f%84%ec%b2%98%eb%a6%ac%eb%9f%89-%ed%94%84%eb%9f%b0%ed%8b%b0%ec%96%b4%eb%a5%bc-%eb%8b%b9%ea%b8%b0%eb%8b%a4" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;&lt;strong&gt;Nemotron 3&lt;/strong&gt; 는 MoE 하이브리드 Mamba–Transformer, &lt;strong&gt;LatentMoE&lt;/strong&gt; , &lt;strong&gt;MTP&lt;/strong&gt; , &lt;strong&gt;NVFP4&lt;/strong&gt; 학습, &lt;strong&gt;멀티-환경 RL&lt;/strong&gt; 을 결합해 “정확도 대비 추론 처리량(accuracy-to-inference-throughput)”을 끌어올리고, &lt;strong&gt;최대 1M tokens 컨텍스트&lt;/strong&gt; 와 &lt;strong&gt;상대 처리량 3.3×&lt;/strong&gt; 를 핵심 메시지로 제시한다 (근거: §Intro/§2.2/§2.3/§2.4/§2.5/§2.6/Fig.2).&lt;/p&gt;</description></item><item><title>[논문리뷰] Pretraining Large Language Models with NVFP4</title><link>https://jaehun.me/posts/%EB%85%BC%EB%AC%B8%EB%A6%AC%EB%B7%B0-pretraining-large-language-models-with-nvfp4/</link><pubDate>Thu, 09 Oct 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/%EB%85%BC%EB%AC%B8%EB%A6%AC%EB%B7%B0-pretraining-large-language-models-with-nvfp4/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2509.25149v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="nvfp4로-4-bit-프리트레이닝을-실전으로-12b를-10t-토큰까지-fp8과-사실상-동급"&gt;NVFP4로 4-bit 프리트레이닝을 실전으로: 12B를 10T 토큰까지, FP8과 사실상 동급&lt;a href="#nvfp4%eb%a1%9c-4-bit-%ed%94%84%eb%a6%ac%ed%8a%b8%eb%a0%88%ec%9d%b4%eb%8b%9d%ec%9d%84-%ec%8b%a4%ec%a0%84%ec%9c%bc%eb%a1%9c-12b%eb%a5%bc-10t-%ed%86%a0%ed%81%b0%ea%b9%8c%ec%a7%80-fp8%ea%b3%bc-%ec%82%ac%ec%8b%a4%ec%83%81-%eb%8f%99%ea%b8%89" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="tldr"&gt;TL;DR&lt;a href="#tldr" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;12B 하이브리드 Mamba-Transformer를 &lt;strong&gt;10T tokens&lt;/strong&gt; 에서 &lt;strong&gt;NVFP4(4-bit)&lt;/strong&gt; 로 프리트레이닝하면 &lt;strong&gt;안정 구간 손실차 &amp;lt;1%&lt;/strong&gt; , 말기 &lt;strong&gt;~1.5%&lt;/strong&gt; 로 &lt;strong&gt;FP8을 근접 추종&lt;/strong&gt; 하고, 다운스트림 성능도 대부분 동급(수학·다국어 일부 +0.9~+3.7pp)이다. &lt;strong&gt;MXFP4 대비 동일 손실에 토큰 +36%&lt;/strong&gt; (1.36T vs 1.0T)가 필요해 &lt;strong&gt;NVFP4의 토큰 효율 우위&lt;/strong&gt; 가 확인된다. (근거: Fig.2, Tab.2, Fig.6)&lt;/p&gt;</description></item><item><title>[논문리뷰] Inference-Time Hyper-Scaling with KV Cache Compression</title><link>https://jaehun.me/posts/%EB%85%BC%EB%AC%B8%EB%A6%AC%EB%B7%B0-inference-time-hyper-scaling-with-kv-cache-compression/</link><pubDate>Tue, 29 Jul 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/%EB%85%BC%EB%AC%B8%EB%A6%AC%EB%B7%B0-inference-time-hyper-scaling-with-kv-cache-compression/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2506.05345v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h2 id="dynamic-memory-sparsificationdms-kv-캐시-8-압축으로-llm-하이퍼-스케일링을-현실로"&gt;Dynamic Memory Sparsification(DMS): KV 캐시 8× 압축으로 LLM 하이퍼-스케일링을 현실로&lt;a href="#dynamic-memory-sparsificationdms-kv-%ec%ba%90%ec%8b%9c-8-%ec%95%95%ec%b6%95%ec%9c%bc%eb%a1%9c-llm-%ed%95%98%ec%9d%b4%ed%8d%bc-%ec%8a%a4%ec%bc%80%ec%9d%bc%eb%a7%81%ec%9d%84-%ed%98%84%ec%8b%a4%eb%a1%9c" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;h2 id="한-줄-요약-tldr"&gt;한 줄 요약 (TL;DR)&lt;a href="#%ed%95%9c-%ec%a4%84-%ec%9a%94%ec%95%bd-tldr" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;&lt;code&gt;1 K&lt;/code&gt; 스텝만의 경량 재적합과 &lt;strong&gt;지연 퇴출 전략&lt;/strong&gt;을 결합한 &lt;strong&gt;DMS&lt;/strong&gt;는 KV 캐시를 최대 &lt;strong&gt;8×&lt;/strong&gt; 압축하면서도 Qwen-R1 32B 기준 AIME 24 &lt;strong&gt;+9.1 pt&lt;/strong&gt; 등 성능을 오히려 끌어올렸다. 결과적으로 동일 연산·메모리 예산에서 &lt;strong&gt;더 길고·더 많은&lt;/strong&gt; 토큰을 실시간으로 생성할 수 있다.&lt;/p&gt;</description></item><item><title>Hogwild! Inference: Parallel LLM Generation via Concurrent Attention</title><link>https://jaehun.me/posts/hogwild-inference-parallel-llm-generation-via-concurrent-attention/</link><pubDate>Thu, 19 Jun 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/hogwild-inference-parallel-llm-generation-via-concurrent-attention/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2504.06261v3"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="hogwild-inference-논문-심층-분석"&gt;Hogwild! Inference 논문 심층 분석&lt;a href="#hogwild-inference-%eb%85%bc%eb%ac%b8-%ec%8b%ac%ec%b8%b5-%eb%b6%84%ec%84%9d" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;제공해주신 &amp;ldquo;Hogwild! Inference: Parallel LLM Generation via Concurrent Attention&amp;rdquo; 논문을 자세히 분석하여 강점과 독창성, 핵심 알고리즘, 그리고 한계점을 설명해 드리겠습니다.&lt;/p&gt;</description></item><item><title>Towards Economical Inference: Enabling DeepSeek’s Multi-Head Latent Attention in Any Transformer-based LLMs</title><link>https://jaehun.me/posts/towards-economical-inference-enabling-deepseeks-multi-head-latent-attention-in-any-transformer-based-llms/</link><pubDate>Mon, 16 Jun 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/towards-economical-inference-enabling-deepseeks-multi-head-latent-attention-in-any-transformer-based-llms/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.14837v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="mha2mla-논문-분석-기존-llm을-경제적으로-만드는-혁신"&gt;MHA2MLA 논문 분석: 기존 LLM을 경제적으로 만드는 혁신&lt;a href="#mha2mla-%eb%85%bc%eb%ac%b8-%eb%b6%84%ec%84%9d-%ea%b8%b0%ec%a1%b4-llm%ec%9d%84-%ea%b2%bd%ec%a0%9c%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%a7%8c%eb%93%9c%eb%8a%94-%ed%98%81%ec%8b%a0" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;제출된 논문 &amp;ldquo;Towards Economical Inference: Enabling DeepSeek&amp;rsquo;s Multi-Head Latent Attention in Any Transformer-based LLMS&amp;quot;는 기존의 대규모 언어 모델(LLM)이 가진 고질적인 문제인 막대한 추론 비용을 해결하기 위한 혁신적이고 실용적인 방법을 제시합니다. 논문의 핵심 내용, 강점, 독창성, 그리고 한계점을 상세히 분석해 드립니다.&lt;/p&gt;</description></item><item><title>X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression</title><link>https://jaehun.me/posts/x-ecomla-upcycling-pre-trained-attention-into-mla-for-efficient-and-extreme-kv-compression/</link><pubDate>Mon, 16 Jun 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/x-ecomla-upcycling-pre-trained-attention-into-mla-for-efficient-and-extreme-kv-compression/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2503.11132v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="x-ecomla-논문-분석-강점-핵심-알고리즘-한계점"&gt;X-EcoMLA 논문 분석: 강점, 핵심 알고리즘, 한계점&lt;a href="#x-ecomla-%eb%85%bc%eb%ac%b8-%eb%b6%84%ec%84%9d-%ea%b0%95%ec%a0%90-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;&amp;ldquo;X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression&amp;rdquo; 논문은 기존의 대규모 언어 모델(LLM)이 가진 메모리 문제를 해결하기 위한 실용적이고 효율적인 접근법을 제시합니다. 본 분석에서는 논문의 독창적인 강점, 예시를 통한 핵심 알고리즘 설명, 그리고 연구의 잠재적 한계점을 심층적으로 다룹니다.&lt;/p&gt;</description></item><item><title>XGRAMMAR: FLEXIBLE AND EFFICIENT STRUCTURED GENERATION ENGINE FOR LARGE LANGUAGE MODELS</title><link>https://jaehun.me/posts/xgrammar-flexible-and-efficient-structured-generation-engine-for-large-language-models/</link><pubDate>Mon, 02 Jun 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/xgrammar-flexible-and-efficient-structured-generation-engine-for-large-language-models/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2411.15100v3"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="xgrammar-논문-상세-분석-유연하고-효율적인-llm-구조화-생성-엔진"&gt;XGrammar 논문 상세 분석: 유연하고 효율적인 LLM 구조화 생성 엔진&lt;a href="#xgrammar-%eb%85%bc%eb%ac%b8-%ec%83%81%ec%84%b8-%eb%b6%84%ec%84%9d-%ec%9c%a0%ec%97%b0%ed%95%98%ea%b3%a0-%ed%9a%a8%ec%9c%a8%ec%a0%81%ec%9d%b8-llm-%ea%b5%ac%ec%a1%b0%ed%99%94-%ec%83%9d%ec%84%b1-%ec%97%94%ec%a7%84" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;제공해주신 논문 &amp;ldquo;XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models&amp;rdquo; [cite: 1]은 대규모 언어 모델(LLM)이 JSON, SQL, 코드 등 특정 구조를 가진 텍스트를 생성해야 하는 요구에 부응하기 위한 새로운 엔진 XGrammar를 제안합니다. 기존의 문맥 자유 문법(Context-Free Grammar, CFG) 기반 제약 디코딩 방식이 가진 실행 시간 오버헤드 문제를 해결하는 데 초점을 맞추고 있습니다. [cite: 2, 3, 4]&lt;/p&gt;</description></item><item><title>RODIMUS*: BREAKING THE ACCURACY-EFFICIENCY TRADE-OFF WITH EFFICIENT ATTENTIONS</title><link>https://jaehun.me/posts/rodimus-breaking-the-accuracy-efficiency-trade-off-with-efficient-attentions/</link><pubDate>Sat, 17 May 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/rodimus-breaking-the-accuracy-efficiency-trade-off-with-efficient-attentions/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2410.06577v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="rodimus-논문-분석-정확도-효율성-균형을-깨는-새로운-언어-모델"&gt;Rodimus* 논문 분석: 정확도-효율성 균형을 깨는 새로운 언어 모델&lt;a href="#rodimus-%eb%85%bc%eb%ac%b8-%eb%b6%84%ec%84%9d-%ec%a0%95%ed%99%95%eb%8f%84-%ed%9a%a8%ec%9c%a8%ec%84%b1-%ea%b7%a0%ed%98%95%ec%9d%84-%ea%b9%a8%eb%8a%94-%ec%83%88%eb%a1%9c%ec%9a%b4-%ec%96%b8%ec%96%b4-%eb%aa%a8%eb%8d%b8" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;대규모 언어 모델(LLM)의 발전은 자연어 처리 분야에 혁신을 가져왔지만, 기존 소프트맥스 어텐션 메커니즘의 높은 연산 비용($\mathcal{O}(T)$ 복잡도)은 효율성 측면에서 한계로 지적되어 왔습니다. [cite: 2] 최근 발표된 &amp;ldquo;RODIMUS*: BREAKING THE ACCURACY-EFFICIENCY TRADE-OFF WITH EFFICIENT ATTENTIONS&amp;rdquo; 논문은 이러한 문제를 해결하기 위해 Rodimus와 Rodimus+라는 새로운 모델을 제시하며, LLM의 정확도를 유지하면서도 연산 복잡도를 획기적으로 낮추는 방법을 제안합니다.&lt;/p&gt;</description></item><item><title>Gemini Embedding: Generalizable Embeddings from Gemini</title><link>https://jaehun.me/posts/gemini-embedding-generalizable-embeddings-from-gemini/</link><pubDate>Mon, 12 May 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/gemini-embedding-generalizable-embeddings-from-gemini/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2503.07891v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="-결론-요약"&gt;✅ 결론 요약&lt;a href="#-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;Gemini Embedding&lt;/strong&gt;은 Google Gemini LLM에서 초기화된 &lt;strong&gt;범용 임베딩 모델로&lt;/strong&gt;, MTEB(Multilingual) 기준 평균 +5.09 점의 성능 향상과 SOTA 달성을 기록하며 &lt;strong&gt;분류, 검색, 클러스터링&lt;/strong&gt; 등 다양한 태스크에서 강력한 일반화 능력을 보여줍니다.&lt;/p&gt;</description></item><item><title>Gemma 3 Technical Report</title><link>https://jaehun.me/posts/gemma-3-technical-report/</link><pubDate>Mon, 12 May 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/gemma-3-technical-report/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2503.19786v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="-결론-요약"&gt;✅ 결론 요약&lt;a href="#-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;&lt;strong&gt;Gemma 3는 장문 컨텍스트(최대 128K), 멀티모달 이미지 입력, 다국어 지원을 갖춘 1B~27B 오픈 모델 계열로, 효율적인 KV 캐시 설계와 향상된 distillation 및 RLHF 기반 후처리로 성능·메모리·활용성에서 매우 뛰어납니다. 특히 27B IT 모델은 Chatbot Arena Elo 1338로 LLaMA 3 70B보다 우위에 있으며, 시각 벤치마크에서도 최고 수준 성능을 달성합니다.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention</title><link>https://jaehun.me/posts/switchhead-accelerating-transformers-with-mixture-of-experts-attention/</link><pubDate>Mon, 14 Apr 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/switchhead-accelerating-transformers-with-mixture-of-experts-attention/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2312.07987v3"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문 「SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention」을 매우 자세하게 읽고 분석한 내용을 바탕으로, 논문의 강점과 독창적인 지점, 핵심 알고리즘의 상세한 설명과 함께 예시 입력을 이용한 동작 과정을 소개하고, 마지막으로 한계점을 명확하게 설명하겠습니다.&lt;/p&gt;</description></item><item><title>SparseTransX: Efficient Training of Translation-Based Knowledge Graph Embeddings Using Sparse Matrix Operations</title><link>https://jaehun.me/posts/sparsetransx-efficient-training-of-translation-based-knowledge-graph-embeddings-using-sparse-matrix-operations/</link><pubDate>Wed, 02 Apr 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/sparsetransx-efficient-training-of-translation-based-knowledge-graph-embeddings-using-sparse-matrix-operations/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.16949"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문 『SparseTransX: Efficient Training of Translation-Based Knowledge Graph Embeddings Using Sparse Matrix Operations』의 강점, 독창적인 지점, 핵심 알고리즘의 예시와 전체적인 과정, 그리고 한계점을 정리하여 전달드립니다.&lt;/p&gt;</description></item><item><title>SELF-DATA DISTILLATION FOR RECOVERING QUALITY IN PRUNED LARGE LANGUAGE MODELS</title><link>https://jaehun.me/posts/self-data-distillation-for-recovering-quality-in-pruned-large-language-models/</link><pubDate>Tue, 25 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/self-data-distillation-for-recovering-quality-in-pruned-large-language-models/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2410.09982"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문의 핵심을 정리하여 결론부터 간략히 제시한 후, 구체적인 수치를 통해 강점 및 독창적인 지점을 설명하고, 논문에서 제안한 핵심 알고리즘을 예시와 함께 설명하며, 논문의 한계점을 논의하겠습니다.&lt;/p&gt;</description></item><item><title>SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention</title><link>https://jaehun.me/posts/sampleattention-near-lossless-acceleration-of-long-context-llm-inference-with-adaptive-structured-sparse-attention/</link><pubDate>Mon, 24 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/sampleattention-near-lossless-acceleration-of-long-context-llm-inference-with-adaptive-structured-sparse-attention/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2406.15486"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="-결론-요약"&gt;✅ 결론 요약&lt;a href="#-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;SampleAttention은 기존 LLM의 attention을 거의 정확도 손실 없이 대체하면서, 최대 2.42배 TTFT(Time-to-First-Token) 지연을 줄이는 구조화된 adaptive sparse attention 기법이다.&lt;/strong&gt;&lt;br&gt;&#10;핵심은 두 가지 sparse 패턴인 &lt;code&gt;local window&lt;/code&gt;와 &lt;code&gt;column stripe&lt;/code&gt;를 활용하여 각 attention head에 대해 동적으로 희소 attention mask를 구성하고, FlashAttention 대비 더 높은 하드웨어 효율성과 가속 성능을 달성한다.&lt;/p&gt;</description></item><item><title>QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving</title><link>https://jaehun.me/posts/qserve-w4a8kv4-quantization-and-system-co-design-for-efficient-llm-serving/</link><pubDate>Tue, 18 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/qserve-w4a8kv4-quantization-and-system-co-design-for-efficient-llm-serving/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2405.04532"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점과-독창적인-지점"&gt;&lt;strong&gt;논문의 강점과 독창적인 지점&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;h4 id="1-강점"&gt;&lt;strong&gt;1. 강점&lt;/strong&gt;&lt;a href="#1-%ea%b0%95%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;저비트 정량화(Quantization)의 실용적 개선:&lt;/strong&gt; 기존 INT4 정량화 기법들이 클라우드 기반 LLM 서빙에서 성능 개선을 보이지 못하는 문제를 해결.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;QoQ(W4A8KV4) 알고리즘 제안:&lt;/strong&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;**4비트 가중치(W4), 8비트 활성화(A8), 4비트 KV 캐시(KV4)**를 적용하여 정량화에 따른 정확도 손실을 최소화.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;진행형 그룹 정량화(Progressive Group Quantization):&lt;/strong&gt; 8비트 중간 표현을 활용하여 INT8 텐서 코어에서 연산을 수행, 기존 INT4 방식보다 높은 성능 제공.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;SmoothAttention 기법:&lt;/strong&gt; 4비트 KV 정량화에 따른 정확도 저하를 완화하는 메커니즘.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;QServe 시스템과 알고리즘 공동 설계(System-Algorithm Co-design):&lt;/strong&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;GPU 서빙 성능 극대화:&lt;/strong&gt; CUDA 코어에서 수행되는 비효율적인 연산을 줄이고, 텐서 코어 활용도를 극대화.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;레지스터 수준 병렬성(Register-Level Parallelism) 활용:&lt;/strong&gt; INT4→INT8 변환 시 감산 후 곱셈(Subtraction after Multiplication) 방식을 적용하여 연산량 감소.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;연산 중심 가중치 재배열(Compute-aware Weight Reordering):&lt;/strong&gt; CUDA 코어에서의 포인터 연산량을 줄여 L1 캐시 활용 최적화.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h4 id="2-독창적인-지점"&gt;&lt;strong&gt;2. 독창적인 지점&lt;/strong&gt;&lt;a href="#2-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;기존 기법&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;QoQ (논문 기법)&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;W4A4의 낮은 정확도 문제&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;W4A8로 INT8 텐서 코어 활용 가능&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;W4A16의 높은 메모리 사용량&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;KV4 도입으로 메모리 효율 개선&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;INT4 GEMM에서 발생하는 CUDA Core 연산 병목&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Register-Level Parallelism으로 해결&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;기존 KV 캐시 정량화의 정확도 저하&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;SmoothAttention으로 키(Key) 값 정규화&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;hr&gt;&#10;&lt;h3 id="핵심-알고리즘-과정-설명"&gt;&lt;strong&gt;핵심 알고리즘 과정 설명&lt;/strong&gt;&lt;a href="#%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ea%b3%bc%ec%a0%95-%ec%84%a4%eb%aa%85" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;h4 id="1-qoq-정량화-알고리즘"&gt;&lt;strong&gt;1. QoQ 정량화 알고리즘&lt;/strong&gt;&lt;a href="#1-qoq-%ec%a0%95%eb%9f%89%ed%99%94-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;p&gt;QoQ는 두 단계의 정량화로 이루어짐:&lt;/p&gt;</description></item><item><title>EFFICIENT LLM INFERENCE USING DYNAMIC INPUT PRUNING AND CACHE-AWARE MASKING</title><link>https://jaehun.me/posts/efficient-llm-inference-using-dynamic-input-pruning-and-cache-aware-masking/</link><pubDate>Wed, 12 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/efficient-llm-inference-using-dynamic-input-pruning-and-cache-aware-masking/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2412.01380"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-핵심-내용과-독창적인-기여"&gt;&lt;strong&gt;논문의 핵심 내용과 독창적인 기여&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ea%b8%b0%ec%97%ac" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문은 &lt;strong&gt;대형 언어 모델(LLM)의 추론 속도를 향상&lt;/strong&gt;시키기 위해 &lt;strong&gt;Dynamic Input Pruning(DIP)&lt;/strong&gt; 및 &lt;strong&gt;Cache-Aware Masking&lt;/strong&gt; 기법을 제안한다. 기존 LLM들은 &lt;strong&gt;메모리 대역폭의 병목 현상&lt;/strong&gt;으로 인해 모바일 디바이스에서 효율적으로 동작하기 어려웠다. 특히 최신 LLM들이 ReLU 대신 SwiGLU를 사용하는데, 이는 자연적인 활성화 희소성이 낮아 기존 동적 희소화 기법이 비효율적이었다.&lt;/p&gt;</description></item><item><title>Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts</title><link>https://jaehun.me/posts/comet-fine-grained-computation-communication-overlapping-for-mixture-of-experts/</link><pubDate>Mon, 10 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/comet-fine-grained-computation-communication-overlapping-for-mixture-of-experts/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.16949"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-핵심-요약-및-기여점"&gt;&lt;strong&gt;논문의 핵심 요약 및 기여점&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%ec%9a%94%ec%95%bd-%eb%b0%8f-%ea%b8%b0%ec%97%ac%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문은 &lt;strong&gt;Comet&lt;/strong&gt;이라는 새로운 MoE(Mixture-of-Experts) 시스템을 제안하여 &lt;strong&gt;계산-통신 오버래핑&lt;/strong&gt;을 더욱 세밀하게 수행함으로써 MoE 모델의 실행 속도를 크게 향상시켰다. 기존 MoE 모델에서 통신 비용이 전체 실행 시간의 47%를 차지하는 문제를 해결하기 위해, &lt;strong&gt;세밀한 수준의(overlapping fine-grained) 계산-통신 오버래핑 기법&lt;/strong&gt;을 도입했다.&lt;/p&gt;</description></item><item><title>LSERVE: EFFICIENT LONG-SEQUENCE LLM SERVING WITH UNIFIED SPARSE ATTENTION</title><link>https://jaehun.me/posts/lserve-efficient-long-sequence-llm-serving-with-unified-sparse-attention/</link><pubDate>Thu, 06 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/lserve-efficient-long-sequence-llm-serving-with-unified-sparse-attention/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.14866"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문의 핵심 내용을 먼저 간략히 요약한 후, 강점과 독창적인 지점을 자세히 설명하고, 핵심 알고리즘의 동작 원리를 예시와 함께 제시한 뒤, 논문의 한계점을 마지막으로 정리하겠습니다.&lt;/p&gt;</description></item><item><title>Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding</title><link>https://jaehun.me/posts/speculate-then-collaborate-fusing-knowledge-of-language-models-during-decoding/</link><pubDate>Tue, 25 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/speculate-then-collaborate-fusing-knowledge-of-language-models-during-decoding/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.08020v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점과-독창적인-지점"&gt;&lt;strong&gt;논문의 강점과 독창적인 지점&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문의 핵심 기여는 &lt;strong&gt;Collaborative Speculative Decoding (CoSD)&lt;/strong&gt; 알고리즘을 제안하여, 훈련 없이도 여러 LLM의 지식을 효과적으로 융합하는 방법을 제시한 점이다. CoSD의 강점은 다음과 같다.&lt;/p&gt;</description></item><item><title>You OnlyPruneOnce: DESIGNING CALIBRATION-FREE MODEL COMPRESSION WITH POLICY LEARNING</title><link>https://jaehun.me/posts/you-onlypruneonce-designing-calibration-free-model-compression-with-policy-learning/</link><pubDate>Tue, 25 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/you-onlypruneonce-designing-calibration-free-model-compression-with-policy-learning/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2501.15296"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-핵심-요약"&gt;&lt;strong&gt;논문의 핵심 요약&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문은 &lt;strong&gt;PruneNet&lt;/strong&gt;이라는 새로운 모델 압축 기법을 제안하며, 기존 방법들의 한계를 극복하고자 한다. 주요 기여점은 다음과 같다:&lt;/p&gt;</description></item><item><title>TypedThinker: Typed Thinking Improves Large Language Model Reasoning</title><link>https://jaehun.me/posts/typedthinker-typed-thinking-improves-large-language-model-reasoning/</link><pubDate>Mon, 24 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/typedthinker-typed-thinking-improves-large-language-model-reasoning/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2410.01952"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="논문의-강점과-독창적인-지점"&gt;논문의 강점과 독창적인 지점&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;h3 id="1-강점"&gt;1. &lt;strong&gt;강점&lt;/strong&gt;&lt;a href="#1-%ea%b0%95%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ul&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;다양한 추론 방식의 활용&lt;/strong&gt;&lt;br&gt;&#10;기존 LLM(대형 언어 모델)의 한계를 극복하기 위해 &lt;strong&gt;연역(deductive), 귀납(inductive), 가설(abductive), 유추(analogical)&lt;/strong&gt; 네 가지의 논리적 추론 방식을 적용함.&lt;br&gt;&#10;각각의 추론 방식이 특정 유형의 문제에서 효과적인지를 실험적으로 분석함.&lt;/p&gt;</description></item><item><title>LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid</title><link>https://jaehun.me/posts/lasp-2-rethinking-sequence-parallelism-for-linear-attention-and-its-hybrid/</link><pubDate>Mon, 17 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/lasp-2-rethinking-sequence-parallelism-for-linear-attention-and-its-hybrid/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.07563v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-및-독창적인-지점"&gt;&lt;strong&gt;논문의 강점 및 독창적인 지점&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;1. 기존 Sequence Parallelism (SP)의 한계 개선&lt;/strong&gt;&lt;br&gt;&#10;기존 SP 기법(LASP-1, Ring Attention 등)은 &lt;code&gt;Right-product-first&lt;/code&gt; 특징을 제대로 활용하지 못하고, 링 스타일(Ring-style) 통신 방식을 사용해 통신-연산 병렬성이 낮았음. LASP-2는 &lt;code&gt;AllGather&lt;/code&gt; 연산을 활용하여 &lt;strong&gt;단일 통신 스텝&lt;/strong&gt; 만으로 전체 메모리 상태를 공유하도록 최적화하여 통신량을 줄이고 계산 병렬성을 극대화함.&lt;/p&gt;</description></item><item><title>Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign</title><link>https://jaehun.me/posts/robust-and-secure-code-watermarking-for-large-language-models-via-ml/crypto-codesign/</link><pubDate>Thu, 13 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/robust-and-secure-code-watermarking-for-large-language-models-via-ml/crypto-codesign/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.02068v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문의 핵심 요약&lt;/p&gt;</description></item><item><title>SmolLM2: When Smol Goes Big Data-Centric Training of a Small Language Mode</title><link>https://jaehun.me/posts/smollm2-when-smol-goes-big-data-centric-training-of-a-small-language-mode/</link><pubDate>Thu, 13 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/smollm2-when-smol-goes-big-data-centric-training-of-a-small-language-mode/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.02737v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문 요약 및 분석: SmolLM2 - When Smol Goes Big&lt;/p&gt;</description></item><item><title>DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning</title><link>https://jaehun.me/posts/deepseek-r1-incentivizing-reasoning-capability-in-llms-via-reinforcement-learning/</link><pubDate>Wed, 12 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/deepseek-r1-incentivizing-reasoning-capability-in-llms-via-reinforcement-learning/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2501.12948v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점과-독창성"&gt;논문의 강점과 독창성&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%84%b1" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;1. 순수 강화학습 기반 모델 개발:&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding</title><link>https://jaehun.me/posts/deepseek-vl2-mixture-of-experts-vision-language-models-for-advanced-multimodal-understanding/</link><pubDate>Wed, 12 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/deepseek-vl2-mixture-of-experts-vision-language-models-for-advanced-multimodal-understanding/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2412.10302v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문-개요-및-강점"&gt;논문 개요 및 강점&lt;a href="#%eb%85%bc%eb%ac%b8-%ea%b0%9c%ec%9a%94-%eb%b0%8f-%ea%b0%95%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;논문 제목:&lt;/strong&gt; &lt;em&gt;DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding&lt;/em&gt;&lt;/p&gt;</description></item><item><title>Humanity's Last Exam</title><link>https://jaehun.me/posts/humanitys-last-exam/</link><pubDate>Wed, 12 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/humanitys-last-exam/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2501.14249v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-및-독창적인-지점"&gt;논문의 강점 및 독창적인 지점&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;문제 인식 및 대응&lt;/strong&gt;: 이 논문은 기존 벤치마크(MMLU 등)가 LLM(대형 언어 모델)에 의해 거의 90% 이상의 정확도를 기록하며 포화 상태에 도달한 문제를 지적하고 있습니다. 이를 해결하기 위해 **Humanity&amp;rsquo;s Last Exam (HLE)**이라는 새로운 벤치마크를 제안합니다.&lt;/p&gt;</description></item><item><title>Qwen2.5-1M Technical Report</title><link>https://jaehun.me/posts/qwen2.5-1m-technical-report/</link><pubDate>Wed, 12 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/qwen2.5-1m-technical-report/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2501.15383v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-개요-및-강점"&gt;논문의 개요 및 강점&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%9c%ec%9a%94-%eb%b0%8f-%ea%b0%95%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;Qwen2.5-1M&lt;/strong&gt;은 Alibaba Group의 Qwen 팀이 개발한 초장문 컨텍스트 지원 대규모 언어 모델 시리즈입니다. 이 모델은 기존 128K 토큰에서 &lt;strong&gt;1백만 토큰&lt;/strong&gt;으로 컨텍스트 길이를 확장했으며, 이를 통해 코드 작성, 문서 요약, 복잡한 질의 응답 등의 작업에서 탁월한 성능을 보여줍니다.&lt;/p&gt;</description></item><item><title>Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at AnyResolution</title><link>https://jaehun.me/posts/qwen2-vl-enhancing-vision-language-models-perception-of-the-world-at-anyresolution/</link><pubDate>Tue, 11 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/qwen2-vl-enhancing-vision-language-models-perception-of-the-world-at-anyresolution/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2409.12191v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문의 강점과 독창적인 지점&lt;/p&gt;</description></item><item><title>DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search</title><link>https://jaehun.me/posts/deepseek-prover-v1.5-harnessing-proof-assistant-feedback-for-reinforcement-learning-and-monte-carlo-tree-search/</link><pubDate>Mon, 10 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/deepseek-prover-v1.5-harnessing-proof-assistant-feedback-for-reinforcement-learning-and-monte-carlo-tree-search/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2408.08152v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-및-독창성"&gt;&lt;strong&gt;논문의 강점 및 독창성&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%84%b1" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;강화 학습과 증명 보조기 피드백의 통합&lt;/strong&gt;&lt;br&gt;&#10;DeepSeek-Prover-V1.5는 증명 보조기(Lean 4)의 피드백을 활용한 강화 학습(RLPAF)을 통해 모델의 성능을 개선했습니다. 기존 모델들은 단순한 감독 학습에 의존했으나, 이 모델은 증명 검증 결과를 직접적으로 학습에 반영함으로써 정확성을 높였습니다.&lt;/p&gt;</description></item></channel></rss>