[논문리뷰] Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus본 논문은 HLA LLM 내부의 계층별 하이브리드화가 Activation Dynamics를 어떻게 재구성하는지 규명하고자 합니다.#Review#Hybrid Linear Attention#Massive Activations#Pre-Attention Spikes#Inter-Spike Plateaus#Attention Sinks#Language Modeling2026년 8월 13일댓글 수 로딩 중
[논문리뷰] SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation본 논문은 기존의 고해상도 Video DiT 모델들이 사용하는 Full-Softmax Attention의 O(N²) 계산 복잡도로 인해 발생하는 심각한 Latency 및 연산 비용 문제를 해결하고자 한다.#Review#Video Generation#Diffusion Transformer#Hybrid Linear Attention#Attention Residuals#Flow Matching#Efficient Inference#Scaling Law2026년 7월 23일댓글 수 로딩 중
[논문리뷰] Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillation본 논문은 하이브리드 모델로의 전환 시 발생하는 부적절한 재귀적 파라미터 초기화 문제를 해결하고자 합니다. 기존 연구들은 Transformer의 가중치를 복사하는 데 집중하지만, 새롭게 도입되는 GDN의 동역학(decay, gate 등)을 고려하지 않아 초기 모델이 최적화되지 않은 상태에서 학습을 시작하게 됩니다 .#Review#Hybrid Linear Attention#Gated DeltaNet#Model Distillation#Initialization#Softmax Attention#Knowledge Distillation#Recurrent Dynamics2026년 6월 18일댓글 수 로딩 중