<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Long-Context on Jaehun's Blog</title><link>https://jaehun.me/categories/long-context/</link><description>Recent content in Long-Context on Jaehun's Blog</description><generator>Hugo</generator><language>ko-kr</language><lastBuildDate>Fri, 11 Sep 2026 09:31:03 +0900</lastBuildDate><atom:link href="https://jaehun.me/categories/long-context/index.xml" rel="self" type="application/rss+xml"/><item><title>Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs</title><link>https://jaehun.me/posts/why-is-video-still-so-expensive-a-survey-of-inference-efficiency-mechanisms-in-video-and-audiovisual-llms/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0900</pubDate><guid>https://jaehun.me/posts/why-is-video-still-so-expensive-a-survey-of-inference-efficiency-mechanisms-in-video-and-audiovisual-llms/</guid><description>&lt;p&gt;&lt;a&#10; href="https://arxiv.org/abs/2609.10355v1"target="_blank"&#10; class="inline-flex items-center gap-1"&#10; &gt;논문 링크&lt;svg class="h-3 w-3 flex-shrink-0" id="external-link" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M15 3h6v6m-11 5L21 3m-3 10v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/&gt;&lt;/svg&gt;&#10; &lt;/a&gt;&lt;/p&gt;&#10;&lt;h2 id="비디오-llm은-왜-아직도-비싼가--추론-효율화-메커니즘-4단계-총정리"&gt;비디오 LLM은 왜 아직도 비싼가? — 추론 효율화 메커니즘 4단계 총정리&lt;a href="#%eb%b9%84%eb%94%94%ec%98%a4-llm%ec%9d%80-%ec%99%9c-%ec%95%84%ec%a7%81%eb%8f%84-%eb%b9%84%ec%8b%bc%ea%b0%80--%ec%b6%94%eb%a1%a0-%ed%9a%a8%ec%9c%a8%ed%99%94-%eb%a9%94%ec%bb%a4%eb%8b%88%ec%a6%98-4%eb%8b%a8%ea%b3%84-%ec%b4%9d%ec%a0%95%eb%a6%ac" class="heading-anchor" aria-label="이 섹션에 대한 링크"&gt;&lt;svg class="h-4 w-4" aria-hidden="true" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;g fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2"&gt;&lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/&gt;&lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — 이 논문은 2022년 말부터 2026년 8월까지 발표된 VideoLLM 추론 효율화 연구 &lt;strong&gt;125편&lt;/strong&gt;을 엔코더–커넥터–LLM 파이프라인의 &lt;strong&gt;4단계&lt;/strong&gt;로 분류하고, &amp;ldquo;같은 호스트 모델 · 같은 입력 프로토콜&amp;quot;에서만 통제 비교를 수행해, 시각 토큰의 약 &lt;strong&gt;25%&lt;/strong&gt; 만 남겨도 정확도가 거의 유지된다는 일관된 결론을 도출한 서베이이다. (근거: Abstract, §I)&lt;/p&gt;</description></item></channel></rss>