[논문리뷰] The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction본 연구는 대규모 MoE 모델을 단일 GPU에서 구동할 때 발생하는 극심한 Memory Wall 문제를 해결하는 데 집중합니다. 기존 연구들은 주로 Weight Quantization이나 Weight Pruning을 통해 모델 크기를 줄여 VRAM 내에 적재하려 시도했으나, 이는 모델의 Accuracy 저하를 초래합니다.#Review#Mixture-of-Experts#SSD#Memory Wall#Routing Prediction#LLM Inference#Throughput2026년 9월 16일댓글 수 로딩 중
[논문리뷰] OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching본 논문은 대규모 언어 모델(LLM)의 long-context 및 agentic workloads가 증가함에 따라 발생하는 HBM의 용량 병목 현상을 해결하고자 한다 .#Review#LLM Inference#KV Cache#Sparse Prefetching#Memory Wall#Speculative Decoding#Lookahead Attention#PD Disaggregation2026년 8월 10일댓글 수 로딩 중