[논문리뷰] On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability본 논문은 대규모 Sparse Mixture-of-Experts (MoE) 모델인 Qwen3.8-Flash-Next를 설계하며, 기존 모델 대비 컴퓨팅 예산을 절감하면서도 동등 이상의 성능을 유지하는 것을 목표로 한다.#Review#Sparse Mixture-of-Experts#Gated DeltaNet#Qwen Sparse Attention#Gated Residual#Muon Optimizer#Long-context Inference2026년 8월 31일댓글 수 로딩 중
[논문리뷰] Dion3: Full-Stack Orthogonal Updates본 논문은 최신 LLM 학습에서 탁월한 성능을 발휘하는 Muon 옵티마이저가 겪는 높은 계산 및 통신 비용 문제를 해결하기 위해 고안되었습니다.#Review#Muon Optimizer#Orthogonalization#Newton-Schulz#Gram Matrix#GPU Kernels#Distributed Training#Low-rank Approximation2026년 8월 16일댓글 수 로딩 중
[논문리뷰] Parallax: Parameterized Local Linear Attention for Language Modeling본 논문은 대규모 언어 모델(LLM) 학습에서 Softmax Attention이 가지는 구조적 한계를 극복하고 효율성을 높이는 것을 목표로 한다.#Review#Local Linear Attention#Language Modeling#Muon Optimizer#Parameterized Attention#Arithmetic Intensity2026년 5월 28일댓글 수 로딩 중
[논문리뷰] Arcee Trinity Large Technical Report본 논문은 희소한 Mixture-of-Experts (MoE) 아키텍처를 기반으로 하는 대규모 언어 모델인 Trinity Large 를 개발하고, 효율적인 학습 및 추론 성능과 높은 안정성을 달성하는 것을 목표로 합니다.#Review#Mixture-of-Experts#Sparse LLM#Training Stability#Load Balancing#MoE#Transformer Architecture#Context Extension#Muon Optimizer2026년 2월 19일댓글 수 로딩 중