[논문리뷰] DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression본 논문은 long-horizon agent 서비스 확산에 따라 발생하는 과도한 prefill 계산 비용 및 대규모 KV cache에 의한 HBM/SSD 용량 병목 현상을 해결하고자 합니다.#Review#Multimodal MoE#KV Cache Compression#Compressed Sparse Attention 2#Causal Encoder-Decoder#SWA Bounded Replay#FP4#Agentic Workloads2026년 9월 17일댓글 수 로딩 중
[논문리뷰] FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution본 논문은 데이터센터 인프라가 아닌 개인용 엣지 기기에서 Frontier-scale MoE 모델을 효율적으로 서비스하기 위한 최적화 기법을 연구합니다.#Review#Edge-Native Inference#MoE#Bandwidth-Adaptive Execution#Agentic Workloads#Resource Management#Memory Hierarchy2026년 8월 18일댓글 수 로딩 중
[논문리뷰] DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference본 논문은 에이전틱 LLM 추론 시 KV-Cache 저장소 I/O가 컴퓨테이션보다 병목 현상을 일으키는 문제를 해결하고자 합니다.#Review#LLM Inference#KV-Cache#Storage Bottleneck#Agentic Workloads#Dual-Path Loading#PD Disaggregation#RDMA#Adaptive Scheduling2026년 2월 25일댓글 수 로딩 중