[논문리뷰] SAS: Simple Attention Sparsification via End-to-End Optimization of Context RankingLong-context inference는 LLM에서 autoregressive generation 시 모든 이전 context token에 대한 dense attention을 요구하며, 이로 인해 cumulative attention cost가 context length에 따라 quadratically 증가하는 심각한 efficiency bottleneck을 초래합니다.#Review#Attention Sparsification#Context Ranking#End-to-End Optimization#Language Modeling Loss#Gated Attention#Long-Context LLMs#Triton Kernel2026년 9월 13일댓글 수 로딩 중
[sglang] AMD GPU에서 FP8 KV 캐시 쓰기 최적화: Triton 커널 융합으로 성능 향상AMD GPU의 FP8 KV 캐시 쓰기 성능을 개선하기 위해 Triton 커널을 융합하여 오버헤드를 줄였습니다.#AMD GPU#FP8#Triton Kernel#KV Cache#Optimization#SGLang2026년 4월 25일댓글 수 로딩 중