BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference논문 링크 모델이 ‘생각을 다시 꺼내 보는’ 순간을 잡아라: BeaconKV가 비콘 쿼리로 추론 KV 캐시를 5.8배 압축하는 법 TL;DR — Large Reasoning Model(LRM)은 긴 … 2026년 9월 7일 15분2609.04971v1