[논문리뷰] SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers본 연구는 Looped Transformer의 성능 향상이 아키텍처 자체의 우수성인지, 아니면 반복 실행으로 인해 증가한 FLOPs와 같은 추가 자원 때문인지 규명하고자 한다.#Review#Looped Transformers#Mixture-of-Experts#Scaling Laws#Compute-Matched#Effective Depth#Inductive Bias#Attention Sink2026년 9월 1일댓글 수 로딩 중
[논문리뷰] Loop the Loopies!본 논문은 Looped Transformer가 고정된 컴퓨팅 자원 내에서 Vanilla Transformer보다 우수한 성능을 낼 수 있도록 하는 compute-matched scaling recipe를 정의합니다.#Review#Looped Transformers#Mixture-of-Experts#Layer-Loop#Compute-Matched Scaling#Post-Training#Reasoning Models2026년 7월 19일댓글 수 로딩 중
[논문리뷰] Parallel Loop Transformer for Efficient Test-Time Computation Scaling본 논문은 Looped Transformer의 고질적인 문제인 순차적인 루프 실행 으로 인한 높은 추론 지연 시간 과 선형적으로 증가하는 KV 캐시 메모리 요구사항 을 해결하는 것을 목표로 합니다.#Review#Large Language Models#Looped Transformers#Inference Efficiency#Parallel Computation#KV Cache Optimization#Gated Sliding-Window Attention#Cross-Loop Parallelism2025년 10월 30일댓글 수 로딩 중