<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Computer Vision on Jaehun's Blog</title><link>https://jaehun.me/tags/computer-vision/</link><description>Recent content in Computer Vision on Jaehun's Blog</description><generator>Hugo</generator><language>ko-kr</language><lastBuildDate>Tue, 08 Sep 2026 13:42:29 +0000</lastBuildDate><atom:link href="https://jaehun.me/tags/computer-vision/index.xml" rel="self" type="application/rss+xml"/><item><title>[논문리뷰]: Radial Attention: O(n log n) Sparse Attention with Energy Decay for Long Video Generation</title><link>https://jaehun.me/posts/%EB%85%BC%EB%AC%B8%EB%A6%AC%EB%B7%B0-radial-attention-on-log-n-sparse-attention-with-energy-decay-for-long-video-generation/</link><pubDate>Mon, 15 Dec 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/%EB%85%BC%EB%AC%B8%EB%A6%AC%EB%B7%B0-radial-attention-on-log-n-sparse-attention-with-energy-decay-for-long-video-generation/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2506.19852v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h2 id="radial-attention-on-log-n-sparse-attention-with-energy-decay로-긴-비디오를-싸게-뽑기"&gt;Radial Attention: O(n log n) Sparse Attention with Energy Decay로 “긴 비디오”를 싸게 뽑기&lt;a href="#radial-attention-on-log-n-sparse-attention-with-energy-decay%eb%a1%9c-%ea%b8%b4-%eb%b9%84%eb%94%94%ec%98%a4%eb%a5%bc-%ec%8b%b8%ea%b2%8c-%eb%bd%91%ea%b8%b0" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;h2 id="한-줄-요약-tldr"&gt;한 줄 요약 (TL;DR)&lt;a href="#%ed%95%9c-%ec%a4%84-%ec%9a%94%ec%95%bd-tldr" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;사후(softmax 이후) attention 에너지가 거리(시간/공간)와 함께 &lt;strong&gt;지수적으로 감쇠&lt;/strong&gt; 한다는 관찰을 바탕으로, 계산 밀도도 같은 방식으로 감쇠시키는 &lt;strong&gt;정적(static) 마스크&lt;/strong&gt;를 설계해 긴 비디오 생성의 훈련/추론 비용을 크게 줄인다. (근거: §4.1–§4.2, Eq.3–Eq.4)&lt;/p&gt;</description></item><item><title>MMInference: Accelerating Pre-filling for Long-Context Visual Language Models via Modality-Aware Permutation Sparse Attention</title><link>https://jaehun.me/posts/mminference-accelerating-pre-filling-for-long-context-visual-language-models-via-modality-aware-permutation-sparse-attention/</link><pubDate>Thu, 19 Jun 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/mminference-accelerating-pre-filling-for-long-context-visual-language-models-via-modality-aware-permutation-sparse-attention/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2504.16083v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="mminference-논문-리뷰-vlm의-긴-컨텍스트-추론-순열로-속도의-벽을-넘다"&gt;MMInference 논문 리뷰: VLM의 긴 컨텍스트 추론, &amp;lsquo;순열&amp;rsquo;로 속도의 벽을 넘다&lt;a href="#mminference-%eb%85%bc%eb%ac%b8-%eb%a6%ac%eb%b7%b0-vlm%ec%9d%98-%ea%b8%b4-%ec%bb%a8%ed%85%8d%ec%8a%a4%ed%8a%b8-%ec%b6%94%eb%a1%a0-%ec%88%9c%ec%97%b4%eb%a1%9c-%ec%86%8d%eb%8f%84%ec%9d%98-%eb%b2%bd%ec%9d%84-%eb%84%98%eb%8b%a4" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;최근 Vision Language Model(VLM)은 이미지와 텍스트를 넘어 긴 비디오까지 이해하는 능력으로 무한한 가능성을 보여주고 있습니다. [cite_start]하지만 수백만 개의 토큰으로 이루어진 긴 비디오를 입력받을 때, 모델이 본격적인 답변 생성을 시작하기 전 입력 전체를 처리하는 &amp;lsquo;Pre-filling&amp;rsquo; 단계에서 엄청난 지연이 발생합니다. [cite: 2] [cite_start]이는 어텐션 메커니즘의 연산량이 입력 길이의 제곱에 비례하여 증가하기 때문인데, 현실적인 서비스 적용에 큰 걸림돌이 되어 왔습니다. [cite: 2, 19]&lt;/p&gt;</description></item><item><title>Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model</title><link>https://jaehun.me/posts/seedream-2.0-a-native-chinese-english-bilingual-image-generation-foundation-model/</link><pubDate>Wed, 16 Apr 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/seedream-2.0-a-native-chinese-english-bilingual-image-generation-foundation-model/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2503.13265v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="-결론-요약-핵심-강점-및-독창성"&gt;✅ 결론 요약 (핵심 강점 및 독창성)&lt;a href="#-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd-%ed%95%b5%ec%8b%ac-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%84%b1" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;Seedream 2.0은 &lt;strong&gt;중국어와 영어 모두를 native하게 이해하고 시각화하는 최초 수준의 이중언어 텍스트-이미지 생성 모델&lt;/strong&gt;이다. 주요 강점은 다음과 같다:&lt;/p&gt;</description></item><item><title>Toward Efficient Inference for Mixture of Experts</title><link>https://jaehun.me/posts/toward-efficient-inference-for-mixture-of-experts/</link><pubDate>Wed, 16 Apr 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/toward-efficient-inference-for-mixture-of-experts/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2503.13265v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="-결론-요약-핵심-기여"&gt;✅ 결론 요약 (핵심 기여)&lt;a href="#-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd-%ed%95%b5%ec%8b%ac-%ea%b8%b0%ec%97%ac" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문 **&amp;ldquo;Toward Efficient Inference for Mixture of Experts&amp;rdquo;**는 Mixture-of-Experts (MoE) 기반 Transformer 모델의 &lt;strong&gt;추론 효율성을 극적으로 향상&lt;/strong&gt;시키는 3가지 핵심 기법을 제안합니다:&lt;/p&gt;</description></item><item><title>On Distributed Larger-Than-Memory Subset Selection With Pairwise Submodular Functions</title><link>https://jaehun.me/posts/on-distributed-larger-than-memory-subset-selection-with-pairwise-submodular-functions/</link><pubDate>Mon, 24 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/on-distributed-larger-than-memory-subset-selection-with-pairwise-submodular-functions/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2402.16442"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="-결론-요약"&gt;✅ 결론 요약&lt;a href="#-%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문은 &amp;ldquo;메모리 용량을 초과하는 대규모 데이터셋에서 대표적인 subset을 효율적으로 선택&amp;quot;하는 문제를 다룬다. 기존 방법들은 중앙 서버가 전체 subset을 메모리에 올릴 수 있어야 한다는 제약이 있었지만, 이 논문은 &lt;strong&gt;중앙 서버 없이도 분산 환경에서 고품질 subset을 선택&lt;/strong&gt;할 수 있는 새로운 알고리즘 2가지를 제안한다:&lt;/p&gt;</description></item><item><title>VOLUT: EFFICIENT VOLUMETRIC STREAMING ENHANCED BY LUT-BASED SUPER-RESOLUTION</title><link>https://jaehun.me/posts/volut-efficient-volumetric-streaming-enhanced-by-lut-based-super-resolution/</link><pubDate>Thu, 06 Mar 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/volut-efficient-volumetric-streaming-enhanced-by-lut-based-super-resolution/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2502.12151"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문 『VoLUT: Efficient Volumetric Streaming Enhanced by LUT-based Super-resolution』을 분석하여 핵심 사항을 아래와 같이 압축하여 전달하고, 강점과 독창성, 알고리즘의 동작 과정, 한계점을 차례로 제시합니다.&lt;/p&gt;</description></item><item><title>Dynamic Diffusion Transformer</title><link>https://jaehun.me/posts/dynamic-diffusion-transformer/</link><pubDate>Tue, 25 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/dynamic-diffusion-transformer/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2410.03456"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h2 id="논문의-핵심-요약"&gt;&lt;strong&gt;논문의 핵심 요약&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;이 논문은 **Dynamic Diffusion Transformer (DyDiT)**라는 새로운 모델을 제안하며, 기존 **Diffusion Transformer (DiT)**의 &lt;strong&gt;과도한 연산량 문제&lt;/strong&gt;를 해결하는 것을 목표로 한다. DyDiT는 &lt;strong&gt;시간축(Timestep)과 공간축(Spatial)에서 동적으로 연산을 조정하는 방식&lt;/strong&gt;을 도입하여 효율성을 크게 향상시켰다.&lt;/p&gt;</description></item><item><title>DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding</title><link>https://jaehun.me/posts/deepseek-vl2-mixture-of-experts-vision-language-models-for-advanced-multimodal-understanding/</link><pubDate>Wed, 12 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/deepseek-vl2-mixture-of-experts-vision-language-models-for-advanced-multimodal-understanding/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2412.10302v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문-개요-및-강점"&gt;논문 개요 및 강점&lt;a href="#%eb%85%bc%eb%ac%b8-%ea%b0%9c%ec%9a%94-%eb%b0%8f-%ea%b0%95%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;논문 제목:&lt;/strong&gt; &lt;em&gt;DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding&lt;/em&gt;&lt;/p&gt;</description></item><item><title>Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at AnyResolution</title><link>https://jaehun.me/posts/qwen2-vl-enhancing-vision-language-models-perception-of-the-world-at-anyresolution/</link><pubDate>Tue, 11 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/qwen2-vl-enhancing-vision-language-models-perception-of-the-world-at-anyresolution/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2409.12191v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문의 강점과 독창적인 지점&lt;/p&gt;</description></item><item><title>Qwen-VL: AVersatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond</title><link>https://jaehun.me/posts/qwen-vl-aversatile-vision-language-model-for-understanding-localization-text-reading-and-beyond/</link><pubDate>Tue, 04 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/qwen-vl-aversatile-vision-language-model-for-understanding-localization-text-reading-and-beyond/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2308.12966v3"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-핵심-요약-및-평가"&gt;&lt;strong&gt;논문의 핵심 요약 및 평가&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%ec%9a%94%ec%95%bd-%eb%b0%8f-%ed%8f%89%ea%b0%80" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;h4 id="1-논문의-강점과-독창적인-지점"&gt;&lt;strong&gt;1. 논문의 강점과 독창적인 지점&lt;/strong&gt;&lt;a href="#1-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;p&gt;이 논문은 &lt;strong&gt;Qwen-VL 시리즈&lt;/strong&gt;라는 대규모 &lt;strong&gt;Vision-Language Model (LVLM)&lt;/strong&gt; 을 소개하며, 기존 모델 대비 &lt;strong&gt;다양한 시각적 이해 및 언어적 응용 능력&lt;/strong&gt;을 향상시킨 것이 핵심이다.&lt;br&gt;&#10;다음과 같은 주요 강점과 차별점이 있다:&lt;/p&gt;</description></item><item><title>Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation</title><link>https://jaehun.me/posts/janusdecouplingvisualencoding-for-unified-multimodal-understanding-and-generation/</link><pubDate>Mon, 03 Feb 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/janusdecouplingvisualencoding-for-unified-multimodal-understanding-and-generation/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2410.13848v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문-요약-및-분석-janus---decoupling-visual-encoding-for-unified-multimodal-understanding-and-generation"&gt;&lt;strong&gt;논문 요약 및 분석: Janus - Decoupling Visual Encoding for Unified Multimodal Understanding and Generation&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8-%ec%9a%94%ec%95%bd-%eb%b0%8f-%eb%b6%84%ec%84%9d-janus---decoupling-visual-encoding-for-unified-multimodal-understanding-and-generation" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;hr&gt;&#10;&lt;h2 id="1-연구의-주요-기여"&gt;&lt;strong&gt;1. 연구의 주요 기여&lt;/strong&gt;&lt;a href="#1-%ec%97%b0%ea%b5%ac%ec%9d%98-%ec%a3%bc%ec%9a%94-%ea%b8%b0%ec%97%ac" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;Janus는 멀티모달 이해(Understanding)와 생성(Generation)을 통합한 자율회귀(Autoregressive) 모델로, 기존 접근 방식의 한계를 해결하기 위해 &lt;strong&gt;시각적 인코딩(Visual Encoding)을 분리하는 전략을 제안&lt;/strong&gt;한다.&lt;/p&gt;</description></item><item><title> SANA: EFFICIENT HIGH-RESOLUTION IMAGE SYN THESIS WITH LINEAR DIFFUSION TRANSFORMERS</title><link>https://jaehun.me/posts/sana-efficient-high-resolution-image-syn-thesis-with-linear-diffusion-transformers/</link><pubDate>Wed, 15 Jan 2025 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/sana-efficient-high-resolution-image-syn-thesis-with-linear-diffusion-transformers/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2410.10629v3"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-독창적-지점-그리고-핵심-내용"&gt;논문의 강점, 독창적 지점, 그리고 핵심 내용&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%8f%85%ec%b0%bd%ec%a0%81-%ec%a7%80%ec%a0%90-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;강점:&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction</title><link>https://jaehun.me/posts/the-truth-is-in-there-improving-reasoning-in-language-models-with-layer-selective-rank-reduction/</link><pubDate>Thu, 26 Dec 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/the-truth-is-in-there-improving-reasoning-in-language-models-with-layer-selective-rank-reduction/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2312.13558"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-및-독창적인-지점"&gt;논문의 강점 및 독창적인 지점&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;핵심 요약&lt;/strong&gt;: 이 논문은 Transformer 기반 대형 언어 모델(LLM)의 특정 층에서 &lt;strong&gt;랭크 감소(LASER)&lt;/strong&gt; 기법을 통해 성능을 개선할 수 있음을 발견했습니다. 특히, 훈련이 끝난 모델에 적용 가능한 후처리 방법으로, 특정 층의 고차 요소(작은 특이값에 해당하는 성분)를 제거하여 성능을 향상시키는 점이 독창적입니다.&lt;/p&gt;</description></item><item><title>The Llama 3 Herd of Models</title><link>https://jaehun.me/posts/the-llama-3-herd-of-models/</link><pubDate>Tue, 24 Dec 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/the-llama-3-herd-of-models/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2407.21783v3"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-독창성-핵심-알고리즘-그리고-한계점-요약"&gt;논문의 강점, 독창성, 핵심 알고리즘, 그리고 한계점 요약&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%8f%85%ec%b0%bd%ec%84%b1-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;hr&gt;&#10;&lt;h4 id="결론-요약"&gt;&lt;strong&gt;결론 요약&lt;/strong&gt;&lt;a href="#%ea%b2%b0%eb%a1%a0-%ec%9a%94%ec%95%bd" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;p&gt;이 논문은 Meta의 최신 언어 모델 &lt;strong&gt;Llama 3&lt;/strong&gt;의 개발, 구조, 학습, 성능 평가, 그리고 안전성을 다룹니다. &lt;strong&gt;405B 매개변수&lt;/strong&gt;를 가지며 &lt;strong&gt;128K 토큰의 긴 문맥 처리&lt;/strong&gt;, 멀티모달 확장, 다국어 지원, 툴 사용을 기본적으로 지원하는 것이 특징입니다. 이 모델은 GPT-4에 필적하거나 특정 분야에서 더 우수한 성능을 보여줍니다. 주요 독창성은 &lt;strong&gt;확장된 학습 스케일과 고품질 데이터 사용&lt;/strong&gt;, &lt;strong&gt;긴 문맥 처리 최적화&lt;/strong&gt;, 그리고 &lt;strong&gt;안전성 보장&lt;/strong&gt;에 있습니다. 하지만, &lt;strong&gt;모델의 높은 자원 요구&lt;/strong&gt;, 일부 언어의 성능 한계, 그리고 안전성 조치의 복잡성이 한계로 지적됩니다.&lt;/p&gt;</description></item><item><title>SageAttention2 Technical Report: Accurate 4 Bit Attention for Plug-and-play Inference Acceleration</title><link>https://jaehun.me/posts/sageattention2-technical-report-accurate-4-bit-attention-for-plug-and-play-inference-acceleration/</link><pubDate>Fri, 20 Dec 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/sageattention2-technical-report-accurate-4-bit-attention-for-plug-and-play-inference-acceleration/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2411.10958v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문의 강점 및 독창성&#10;1.&#9;효율성 증가:&#10;•&#9;SageAttention2는 4비트 정밀도(INT4)와 FP8(8비트 부동소수점)을 활용하여 기존 FlashAttention2 대비 약 3.1배 빠른 연산 속도를 제공함.&#10;•&#9;RTX4090에서 최대 485 TOPS 성능을 달성하며, FlashAttention2와 xformers 대비 각각 3.1배, 5.4배 속도 향상을 입증.&#10;2.&#9;정확도 유지:&#10;•&#9;INT4 및 FP8로 매트릭스를 양자화하면서도 End-to-End 정확도 손실이 미미함.&#10;•&#9;텍스트, 이미지, 비디오 생성 모델에서 정확도 손실 없이 기존 모델의 성능을 유지.&#10;3.&#9;적응형 양자화 기법:&#10;•&#9;특정 레이어와 타임스텝에서 높은 정확도를 유지하기 위해 INT8 양자화를 혼합하는 방법을 제안.&#10;•&#9;이를 통해 다양한 입력과 모델 구조에서도 높은 범용성을 보장.&#10;4.&#9;전처리 개선:&#10;•&#9;Smooth Q와 Smooth V 방법론을 통해 양자화 오류를 줄여 INT4의 제한된 수치 범위 내에서도 정확도를 향상.&lt;/p&gt;</description></item><item><title>Token Merging: Your ViT But Faster</title><link>https://jaehun.me/posts/token-merging-your-vit-but-faster/</link><pubDate>Fri, 13 Dec 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/token-merging-your-vit-but-faster/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2210.09461"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;논문의 주요 내용과 분석을 다음과 같이 요약하겠습니다:&lt;/p&gt;</description></item><item><title>DeepCache: Accelerating Diffusion Models for Free</title><link>https://jaehun.me/posts/deepcache-accelerating-diffusion-models-for-free/</link><pubDate>Mon, 09 Dec 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/deepcache-accelerating-diffusion-models-for-free/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2312.00858v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점과-독창성"&gt;논문의 강점과 독창성&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%84%b1" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;강점&lt;/strong&gt;:&lt;/p&gt;</description></item><item><title>LOOK-M:Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference</title><link>https://jaehun.me/posts/look-mlook-once-optimization-in-kv-cache-for-efficient-multimodal-long-context-inference/</link><pubDate>Wed, 27 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/look-mlook-once-optimization-in-kv-cache-for-efficient-multimodal-long-context-inference/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2406.18139"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-및-독창성"&gt;논문의 강점 및 독창성&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%84%b1" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;ol&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;문제 정의의 명확성&lt;/strong&gt;:&lt;/p&gt;</description></item><item><title>SPARSEVLM: VISUAL TOKEN SPARSIFICATION FOR EFFICIENT VISION-LANGUAGE MODEL INFERENCE</title><link>https://jaehun.me/posts/sparsevlm-visual-token-sparsification-for-efficient-vision-language-model-inference/</link><pubDate>Wed, 20 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/sparsevlm-visual-token-sparsification-for-efficient-vision-language-model-inference/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2410.04417"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="강점-및-독창적인-지점"&gt;강점 및 독창적인 지점&lt;a href="#%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;논문의 강점:&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration</title><link>https://jaehun.me/posts/vl-cache-sparsity-and-modality-aware-kv-cache-compression-for-vision-language-model-inference-acceleration/</link><pubDate>Mon, 18 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/vl-cache-sparsity-and-modality-aware-kv-cache-compression-for-vision-language-model-inference-acceleration/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2410.23317"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="q--이-논문을-아주-자세하게-읽고-논문의-강점과-독창적인-지점을-설명해주고-핵심-알고리즘을-예시-입력을-들어서-전체적인-과정을-설명해줘-추가적으로-논문의-한계점에-대해서도-알려줘"&gt;Q : 이 논문을 아주 자세하게 읽고 논문의 강점과 독창적인 지점을 설명해주고 핵심 알고리즘을 예시 입력을 들어서 전체적인 과정을 설명해줘 추가적으로 논문의 한계점에 대해서도 알려줘&lt;a href="#q--%ec%9d%b4-%eb%85%bc%eb%ac%b8%ec%9d%84-%ec%95%84%ec%a3%bc-%ec%9e%90%ec%84%b8%ed%95%98%ea%b2%8c-%ec%9d%bd%ea%b3%a0-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90%ea%b3%bc-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a3%bc%ea%b3%a0-%ed%95%b5%ec%8b%ac-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98%ec%9d%84-%ec%98%88%ec%8b%9c-%ec%9e%85%eb%a0%a5%ec%9d%84-%eb%93%a4%ec%96%b4%ec%84%9c-%ec%a0%84%ec%b2%b4%ec%a0%81%ec%9d%b8-%ea%b3%bc%ec%a0%95%ec%9d%84-%ec%84%a4%eb%aa%85%ed%95%b4%ec%a4%98-%ec%b6%94%ea%b0%80%ec%a0%81%ec%9c%bc%eb%a1%9c-%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%9c%ea%b3%84%ec%a0%90%ec%97%90-%eb%8c%80%ed%95%b4%ec%84%9c%eb%8f%84-%ec%95%8c%eb%a0%a4%ec%a4%98" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문-요약-및-분석"&gt;논문 요약 및 분석&lt;a href="#%eb%85%bc%eb%ac%b8-%ec%9a%94%ec%95%bd-%eb%b0%8f-%eb%b6%84%ec%84%9d" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;논문 제목:&lt;/strong&gt; VL-CACHE: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration&lt;/p&gt;</description></item><item><title>Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models</title><link>https://jaehun.me/posts/deep-compression-autoencoder-for-efficient-high-resolution-diffusion-models/</link><pubDate>Thu, 14 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/deep-compression-autoencoder-for-efficient-high-resolution-diffusion-models/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2410.10733v2"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2410.10733v2&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;아래 글은 Scholar GPT와 대화를 나눈 내용입니다.&lt;/p&gt;&#10;&lt;h1 id="q--논문의-핵심-내용과-강점-알고리즘-설명-그리고-한계점"&gt;Q : 논문의 핵심 내용과 강점, 알고리즘 설명, 그리고 한계점&lt;a href="#q--%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%ea%b0%95%ec%a0%90-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문-분석-및-요약-deep-compression-autoencoder-for-efficient-high-resolution-diffusion-models"&gt;논문 분석 및 요약: &amp;ldquo;Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models&amp;rdquo;&lt;a href="#%eb%85%bc%eb%ac%b8-%eb%b6%84%ec%84%9d-%eb%b0%8f-%ec%9a%94%ec%95%bd-deep-compression-autoencoder-for-efficient-high-resolution-diffusion-models" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;h4 id="1-논문의-개요-및-주요-내용"&gt;1. &lt;strong&gt;논문의 개요 및 주요 내용&lt;/strong&gt;&lt;a href="#1-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%9c%ec%9a%94-%eb%b0%8f-%ec%a3%bc%ec%9a%94-%eb%82%b4%ec%9a%a9" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;p&gt;이 논문은 고해상도 이미지 생성을 위한 효율적인 오토인코더인 **Deep Compression Autoencoder (DC-AE)**를 제안합니다. 기존의 잠재 확산 모델(latent diffusion models)은 오토인코더를 활용하여 고해상도 이미지를 잠재 공간으로 압축하여 계산 비용을 줄이지만, 공간 압축 비율이 높아질수록(예: 64배, 128배) 재구성 정확도가 크게 떨어지는 문제가 있었습니다. DC-AE는 이러한 문제를 해결하기 위해 &lt;strong&gt;Residual Autoencoding&lt;/strong&gt;과 &lt;strong&gt;Decoupled High-Resolution Adaptation&lt;/strong&gt;이라는 두 가지 핵심 기술을 도입했습니다.&lt;/p&gt;</description></item><item><title>HART Efficient Visual Generation with Hybrid Autoregressive Transformer</title><link>https://jaehun.me/posts/hart-efficient-visual-generation-with-hybrid-autoregressive-transformer/</link><pubDate>Thu, 14 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/hart-efficient-visual-generation-with-hybrid-autoregressive-transformer/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2410.10812"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2410.10812&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;아래 글은 Scholar GPT와 대화를 나눈 내용입니다.&lt;/p&gt;&#10;&lt;h1 id="q--논문의-핵심-내용과-강점-알고리즘-설명-그리고-한계점"&gt;Q : 논문의 핵심 내용과 강점, 알고리즘 설명, 그리고 한계점&lt;a href="#q--%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%ea%b0%95%ec%a0%90-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-강점-및-독창적인-지점"&gt;논문의 강점 및 독창적인 지점&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%95%ec%a0%90-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%a0%81%ec%9d%b8-%ec%a7%80%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;이 논문은 Hybrid Autoregressive Transformer (HART)라는 새로운 이미지 생성 모델을 제안하며, 이 모델은 특히 높은 효율성을 자랑합니다. HART는 1024x1024 해상도의 이미지를 직접 생성할 수 있으며, 기존의 확산 모델과 비교하여 다음과 같은 독창적인 강점이 있습니다:&lt;/p&gt;</description></item><item><title>Learning Transferable Visual Models From Natural Language Supervision</title><link>https://jaehun.me/posts/learning-transferable-visual-models-from-natural-language-supervision/</link><pubDate>Thu, 14 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/learning-transferable-visual-models-from-natural-language-supervision/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2103.00020"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2103.00020&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;아래 글은 Scholar GPT와 대화를 나눈 내용입니다.&lt;/p&gt;&#10;&lt;h1 id="q--논문의-핵심-내용과-강점-알고리즘-설명-그리고-한계점"&gt;Q : 논문의 핵심 내용과 강점, 알고리즘 설명, 그리고 한계점&lt;a href="#q--%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%ea%b0%95%ec%a0%90-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;p&gt;이 논문은 &amp;ldquo;Learning Transferable Visual Models From Natural Language Supervision&amp;quot;이라는 제목을 가진 CLIP (Contrastive Language-Image Pre-training) 모델에 대한 연구입니다.&lt;/p&gt;</description></item><item><title>Condition-Aware Neural Network for Controlled Image Generation</title><link>https://jaehun.me/posts/condition-aware-neural-network-for-controlled-image-generation/</link><pubDate>Wed, 13 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/condition-aware-neural-network-for-controlled-image-generation/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2404.01143"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2404.01143&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;아래글은 Scholar GPT와 대화를 나눈 내용입니다.&lt;/p&gt;&#10;&lt;h1 id="q--논문의-핵심-내용과-강점-알고리즘-설명-그리고-한계점"&gt;Q : 논문의 핵심 내용과 강점, 알고리즘 설명, 그리고 한계점&lt;a href="#q--%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%ea%b0%95%ec%a0%90-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-요약-및-분석"&gt;논문의 요약 및 분석&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ec%9a%94%ec%95%bd-%eb%b0%8f-%eb%b6%84%ec%84%9d" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;h4 id="1-논문의-개요"&gt;1. 논문의 개요&lt;a href="#1-%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%9c%ec%9a%94" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;p&gt;논문 **&amp;ldquo;Condition-Aware Neural Network (CAN)&amp;rdquo;**은 이미지 생성 모델에 제어 기능을 추가하기 위해 새로운 접근 방식을 제안합니다. 이 모델은 기존의 조건 제어 방법과는 달리, &lt;strong&gt;조건에 따라 동적으로 네트워크의 가중치를 조정&lt;/strong&gt;하여 이미지 생성 과정을 제어합니다. 이 방식은 기존의 GAN, Transformer 기반의 이미지 생성 모델에서 주로 사용되던 &lt;strong&gt;특징 공간 조작 대신 가중치 공간을 조작&lt;/strong&gt;하는 것을 목표로 합니다.&lt;/p&gt;</description></item><item><title>DistriFusion Distributed Parallel Inference for High-Resolution Diffusion Models</title><link>https://jaehun.me/posts/distrifusion-distributed-parallel-inference-for-high-resolution-diffusion-models/</link><pubDate>Wed, 13 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/distrifusion-distributed-parallel-inference-for-high-resolution-diffusion-models/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2402.19481"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2402.19481&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;아래글은 Scholar GPT와 대화를 나눈 내용입니다.&lt;/p&gt;&#10;&lt;h1 id="q--논문의-핵심-내용과-강점-알고리즘-설명-그리고-한계점"&gt;Q : 논문의 핵심 내용과 강점, 알고리즘 설명, 그리고 한계점&lt;a href="#q--%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%ea%b0%95%ec%a0%90-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문의-개요-및-독창성"&gt;논문의 개요 및 독창성&lt;a href="#%eb%85%bc%eb%ac%b8%ec%9d%98-%ea%b0%9c%ec%9a%94-%eb%b0%8f-%eb%8f%85%ec%b0%bd%ec%84%b1" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;논문 &amp;ldquo;DistriFusion: Distributed Parallel Inference for High-Resolution Diffusion Models&amp;quot;은 고해상도 이미지 생성에서 발생하는 연산 병목을 해결하기 위해 다중 GPU를 활용하여 Diffusion 모델의 추론 속도를 대폭 개선하는 알고리즘을 제안합니다. 기존의 Diffusion 모델은 대규모의 연산량으로 인해 고해상도 이미지 생성 시 실시간 응용에 적합하지 않았습니다. DistriFusion은 고해상도 이미지를 여러 패치(patch)로 나누어 각 GPU에서 병렬로 처리하면서도, 패치 간 상호작용 문제를 해결하기 위해 비동기 통신과 이전 단계의 특징 맵을 재활용하는 접근 방식을 도입합니다.&lt;/p&gt;</description></item><item><title>FastComposer Tuning-Free Multi-Subject Image Generation with Localized Attention</title><link>https://jaehun.me/posts/fastcomposer-tuning-free-multi-subject-image-generation-with-localized-attention/</link><pubDate>Wed, 13 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/fastcomposer-tuning-free-multi-subject-image-generation-with-localized-attention/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2305.10431"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2305.10431&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;아래글은 Scholar GPT와 대화를 나눈 내용입니다.&lt;/p&gt;&#10;&lt;h1 id="q--논문의-핵심-내용과-강점-알고리즘-설명-그리고-한계점"&gt;Q : 논문의 핵심 내용과 강점, 알고리즘 설명, 그리고 한계점&lt;a href="#q--%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%ea%b0%95%ec%a0%90-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문-요약-및-주요-내용-분석"&gt;논문 요약 및 주요 내용 분석&lt;a href="#%eb%85%bc%eb%ac%b8-%ec%9a%94%ec%95%bd-%eb%b0%8f-%ec%a3%bc%ec%9a%94-%eb%82%b4%ec%9a%a9-%eb%b6%84%ec%84%9d" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;논문 제목&lt;/strong&gt;: &amp;ldquo;FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention&amp;rdquo;&lt;br&gt;&#10;&lt;strong&gt;저자&lt;/strong&gt;: Guangxuan Xiao, Tianwei Yin, William T. Freeman, Frédo Durand, Song Han (MIT)&lt;/p&gt;</description></item><item><title>VILA On Pre-training for Visual Language Models</title><link>https://jaehun.me/posts/vila-on-pre-training-for-visual-language-models/</link><pubDate>Wed, 13 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/vila-on-pre-training-for-visual-language-models/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2312.07533"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2312.07533&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;아래글은 Scholar GPT와 대화를 나눈 내용입니다.&lt;/p&gt;&#10;&lt;h1 id="q--논문의-핵심-내용과-강점-알고리즘-설명-그리고-한계점"&gt;Q : 논문의 핵심 내용과 강점, 알고리즘 설명, 그리고 한계점&lt;a href="#q--%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%ea%b0%95%ec%a0%90-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문-분석-vila-on-pre-training-for-visual-language-models"&gt;논문 분석: &lt;strong&gt;VILA: On Pre-training for Visual Language Models&lt;/strong&gt;&lt;a href="#%eb%85%bc%eb%ac%b8-%eb%b6%84%ec%84%9d-vila-on-pre-training-for-visual-language-models" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;h4 id="1-논문의-배경-및-목표"&gt;&lt;strong&gt;1. 논문의 배경 및 목표&lt;/strong&gt;&lt;a href="#1-%eb%85%bc%eb%ac%b8%ec%9d%98-%eb%b0%b0%ea%b2%bd-%eb%b0%8f-%eb%aa%a9%ed%91%9c" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h4&gt;&lt;p&gt;이 논문은 **Visual Language Model (VLM)**을 개선하기 위한 &lt;strong&gt;사전 학습(pre-training)&lt;/strong&gt; 기법을 연구합니다. 최근 **대규모 언어 모델(LLM)**의 성공을 기반으로, 시각적 입력을 처리할 수 있는 &lt;strong&gt;멀티모달 모델&lt;/strong&gt;이 주목받고 있습니다. 그러나 기존 연구는 주로 **시각적 언어 지시 조정(visual instruction tuning)**에 집중되어 있었으며, &lt;strong&gt;사전 학습 단계&lt;/strong&gt;에 대한 심층적인 연구는 부족했습니다.&lt;/p&gt;</description></item><item><title>VILA-U a Unified Foundation Model Integrating Visual Understanding and Generation</title><link>https://jaehun.me/posts/vila-u-a-unified-foundation-model-integrating-visual-understanding-and-generation/</link><pubDate>Wed, 13 Nov 2024 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/vila-u-a-unified-foundation-model-integrating-visual-understanding-and-generation/</guid><description>&lt;p&gt;논문 : &lt;a&#10; href="https://arxiv.org/abs/2409.04429"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;https://arxiv.org/abs/2409.04429&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;아래글은 Scholar GPT와 대화를 나눈 내용입니다.&lt;/p&gt;&#10;&lt;h1 id="q--논문의-핵심-내용과-강점-알고리즘-설명-그리고-한계점"&gt;Q : 논문의 핵심 내용과 강점, 알고리즘 설명, 그리고 한계점&lt;a href="#q--%eb%85%bc%eb%ac%b8%ec%9d%98-%ed%95%b5%ec%8b%ac-%eb%82%b4%ec%9a%a9%ea%b3%bc-%ea%b0%95%ec%a0%90-%ec%95%8c%ea%b3%a0%eb%a6%ac%ec%a6%98-%ec%84%a4%eb%aa%85-%ea%b7%b8%eb%a6%ac%ea%b3%a0-%ed%95%9c%ea%b3%84%ec%a0%90" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h1&gt;&lt;h3 id="논문-요약-및-강점-분석"&gt;논문 요약 및 강점 분석&lt;a href="#%eb%85%bc%eb%ac%b8-%ec%9a%94%ec%95%bd-%eb%b0%8f-%ea%b0%95%ec%a0%90-%eb%b6%84%ec%84%9d" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h3&gt;&lt;p&gt;논문 제목: &lt;strong&gt;VILA-U: A Unified Foundation Model Integrating Visual Understanding and Generation&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>