Large Language Models

11페이지

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

논문 링크 X-AuT: 행동 프로브와 크로스스케일 증류로 음성 LLM의 오디오 인코더를 깎아내는 방법한 줄 요약 (TL;DR)음성 LLM의 오디오 인코더(오디오 Transformer)를 18→16→14 레이어로 점진적으로 프루닝 하면서, 짧은 행동 프 …

12분
Efficient Inference Large Language Models Transformer Machine Learning Knowledge Distillation Model Compression

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

논문 링크 비디오 LLM은 왜 아직도 비싼가? — 추론 효율화 메커니즘 4단계 총정리TL;DR — 이 논문은 2022년 말부터 2026년 8월까지 발표된 VideoLLM 추론 효율화 연구 125편을 엔코더–커넥터–LLM 파이프라인의 4단계로 분류하고, …

14분
Long Context KV Cache Multimodal Learning Vision-Language Model Efficient Inference Large Language Models