[논문리뷰] Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation본 논문은 대규모 Transformer 모델의 파라미터 수 증가가 메모리 및 학습 비용의 병목으로 이어지는 문제를 해결하고자 합니다. 기존의 표준 Transformer는 깊이를 늘릴 때마다 새로운 가중치 텐서를 생성하여 모델의 파라미터 footprint를 기하급수적으로 증가시킵니다.#Review#Transformer#Recurrent Depth#Weight Sharing#Gated Recurrent Transformer#Memory Efficiency#Model Scaling#Adaptive Computation2026년 8월 26일댓글 수 로딩 중