게시글

기술과 일상에 대한 글 목록입니다.

546페이지
7 / 91

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

논문 링크 X-AuT: 행동 프로브와 크로스스케일 증류로 음성 LLM의 오디오 인코더를 깎아내는 방법한 줄 요약 (TL;DR)음성 LLM의 오디오 인코더(오디오 Transformer)를 18→16→14 레이어로 점진적으로 프루닝 하면서, 짧은 행동 프 …

12분
Efficient Inference Large Language Models Transformer Machine Learning Knowledge Distillation Model Compression

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

논문 링크 비디오 LLM은 왜 아직도 비싼가? — 추론 효율화 메커니즘 4단계 총정리TL;DR — 이 논문은 2022년 말부터 2026년 8월까지 발표된 VideoLLM 추론 효율화 연구 125편을 엔코더–커넥터–LLM 파이프라인의 4단계로 분류하고, …

14분
Long Context KV Cache Multimodal Learning Vision-Language Model Efficient Inference Large Language Models

Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training

논문 링크 대규모·장문맥 RL 사후학습에서 Speculative Decoding을 실용화하는 두 가지 설계TL;DR : RL 사후학습의 벽-클록 시간을 지배하는 롤아웃 생성을 speculative decoding으로 가속하고, 여기에 온라인 드래프트 …

11분
Speculative Decoding Long Context Reinforcement Learning Distributed Computing Efficient Training Inference Acceleration