[논문리뷰] ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding본 논문은 실시간 영상 스트리밍 환경에서 MLLM이 겪는 과도한 연산 비용과 효율성 저하 문제를 해결합니다. 기존의 많은 연구는 모든 프레임을 모델의 전체 레이어(Full-depth)로 미리 처리(Prefill)하여 KV Cache를 생성하는데, 이는 매우 비용이 많이 들고 불필요한 연산을 유발합니다 .#Review#Streaming Video Understanding#MLLM#KV Cache#Efficient Inference#Query-Agnostic#Selective Retrieval2026년 9월 6일댓글 수 로딩 중
[논문리뷰] RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction본 논문은 query-agnostic KV 캐시 압축 환경에서 발생하는 성능 급락 문제를 해결하고자 합니다. 기존 연구들은 주로 중요한 KV 쌍을 선별(Selection)하는 데 집중했으나, 압축률이 높아질수록 선별된 정보만으로는 모델 성능을 유지하기 어렵습니다.#Review#KV Cache Eviction#Query-Agnostic#LoRA#Self-Distillation#Long-Context LLM#Context Reconstruction2026년 8월 4일댓글 수 로딩 중