[논문리뷰] FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving본 논문은 Long-context LLM의 prefilling 단계에서 발생하는 self-attention의 quadratic complexity 병목 문제를 해결하기 위해 FlashPrefill V2를 제안합니다.#Review#Long-context LLM#Sparse Attention#FlashAttention#FP8 Inference#KV Cache#LLM Serving2026년 8월 20일댓글 수 로딩 중