[논문리뷰] Fast Weight Attention for Continual Learning본 논문은 Transformer의 KV cache가 시퀀스 길이에 따라 O(N²)의 비용을 소모하며 발생하는 비효율성 및 Continual Learning 환경에서의 파괴적 망각(catastrophic interference) 문제를 해결하고자 합니다.#Review#Continual Learning#Fast Weight#Linear Attention#State Space Models#Online Gradient Descent#Ridge Regression2026년 8월 30일댓글 수 로딩 중