[논문리뷰] On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability본 논문은 대규모 Sparse Mixture-of-Experts (MoE) 모델인 Qwen3.8-Flash-Next를 설계하며, 기존 모델 대비 컴퓨팅 예산을 절감하면서도 동등 이상의 성능을 유지하는 것을 목표로 한다.#Review#Sparse Mixture-of-Experts#Gated DeltaNet#Qwen Sparse Attention#Gated Residual#Muon Optimizer#Long-context Inference2026년 8월 31일댓글 수 로딩 중
[논문리뷰] HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts본 논문은 기존의 MoE 압축 방식들이 전문가 간의 결합 가능성을 평가할 때 사용하는 pairwise 점수의 구조적 한계를 해결하고자 합니다.#Review#Sparse Mixture-of-Experts#Simplicial Complex#Hodge Decomposition#Harmonic Kernel#Model Compression#Topological Deep Learning2026년 5월 17일댓글 수 로딩 중