[논문리뷰] TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration본 논문은 LLM의 Long-context prefill 단계에서 발생하는 쿼리-키 간 Quadratic 연산과 메모리 병목 문제를 해결하기 위해 TileMix를 제안한다.#Review#LLM Inference#Mixed-Precision Attention#Tile-Centric#Online Softmax#Quantization#Long-Context#Kernel Optimization2026년 8월 24일댓글 수 로딩 중
[논문리뷰] Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification본 논문은 대규모 비디오 생성 모델에서 긴 토큰 시퀀스로 인해 발생하는 quadratic complexity의 self-attention 병목 현상을 해결하고자 합니다.#Review#Diffusion Transformers#Video Generation#Attention Sparsification#Online Softmax#Training-free#On-the-fly#Proxy-score Reuse2026년 7월 27일댓글 수 로딩 중