[논문리뷰] Grouped Value Attention: Efficient KV Caching via On-Demand Key ReconstructionThe paper is 'Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction' by Vishesh Tripathi, Abhay Kumar, and Ramsha Khan.#Review#KV Cache#Grouped Value Attention#Autoregressive Decoding#Transformer#Memory Efficiency#Decoupled RoPE#Linear Reconstruction2026년 9월 14일댓글 수 로딩 중