[논문리뷰] DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression본 논문은 long-horizon agent 서비스 확산에 따라 발생하는 과도한 prefill 계산 비용 및 대규모 KV cache에 의한 HBM/SSD 용량 병목 현상을 해결하고자 합니다.#Review#Multimodal MoE#KV Cache Compression#Compressed Sparse Attention 2#Causal Encoder-Decoder#SWA Bounded Replay#FP4#Agentic Workloads2026년 9월 17일댓글 수 로딩 중