[논문리뷰] Rethinking the Role of Efficient Attention in Hybrid Architectures본 논문은 하이브리드 아키텍처에서 Efficient Attention이 모델의 장거리 문맥 학습 능력에 미치는 영향을 체계적으로 규명하는 것을 목표로 합니다.#Review#Hybrid Architecture#Efficient Attention#Full Attention#Scaling Law#Long-Context Capability#Optimization Prior#Large-Window Laziness2026년 6월 16일댓글 수 로딩 중
[논문리뷰] SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space대규모 언어 모델(LLM)에서 quadratic 연산 복잡성 을 갖는 full attention 의 한계를 극복하기 위해, sparse attention 의 성능 저하 및 부족한 sparsity 문제를 해결하고자 합니다.#Review#Sparse Attention#Full Attention#Large Language Models (LLMs)#Context Length#Attention Sparsity#Alignment Loss#Long-Context Extrapolation2025년 11월 25일댓글 수 로딩 중