[논문리뷰] On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability본 논문은 대규모 Sparse Mixture-of-Experts (MoE) 모델인 Qwen3.8-Flash-Next를 설계하며, 기존 모델 대비 컴퓨팅 예산을 절감하면서도 동등 이상의 성능을 유지하는 것을 목표로 한다.#Review#Sparse Mixture-of-Experts#Gated DeltaNet#Qwen Sparse Attention#Gated Residual#Muon Optimizer#Long-context Inference2026년 8월 31일댓글 수 로딩 중
[논문리뷰] Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference본 논문은 기존 long-context LLM 추론에서 발생하는 quadratic computational complexity와 기존 하이브리드 어텐션 기법들의 한계를 해결하고자 합니다.#Review#Large Language Models#Long-context Inference#Hybrid Attention#Dynamic Routing#Layer-level Sparsity#Context-aware2026년 4월 9일댓글 수 로딩 중