[논문리뷰] OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching본 논문은 대규모 언어 모델(LLM)의 long-context 및 agentic workloads가 증가함에 따라 발생하는 HBM의 용량 병목 현상을 해결하고자 한다 .#Review#LLM Inference#KV Cache#Sparse Prefetching#Memory Wall#Speculative Decoding#Lookahead Attention#PD Disaggregation2026년 8월 10일댓글 수 로딩 중
[논문리뷰] DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference본 논문은 에이전틱 LLM 추론 시 KV-Cache 저장소 I/O가 컴퓨테이션보다 병목 현상을 일으키는 문제를 해결하고자 합니다.#Review#LLM Inference#KV-Cache#Storage Bottleneck#Agentic Workloads#Dual-Path Loading#PD Disaggregation#RDMA#Adaptive Scheduling2026년 2월 25일댓글 수 로딩 중