[논문리뷰] MassAlloc Attention: Let Attention Allocate Its Own Compute본 논문은 Long-context Full softmax attention이 상당한 computation 및 memory traffic을 유발하는 문제에 주목합니다.#Review#MassAlloc Attention#Sparse Attention#Compute Allocation#Online Softmax#Long-Context LLMs#Transformer Efficiency#Perplexity#Latency2026년 9월 28일댓글 수 로딩 중