[논문리뷰] Expert-Space Exploration in MoE Reinforcement Learning본 논문은 기존의 LLM post-training 연구가 token-level의 sampling에만 집중하고, MoE 모델의 핵심인 expert routing의 다양성 활용에는 소홀했다는 점을 지적하며 ESRL을 제안합니다.#Review#MoE#Reinforcement Learning#Expert Routing#Exploration#Rollout Routing Replay2026년 9월 14일댓글 수 로딩 중
[논문리뷰] Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers본 논문은 Mixture-of-Experts (MoE) 모델 의 강화 학습(RL) 훈련 과정에서 발생하는 불안정성, 특히 훈련-추론 간 라우팅 동작의 불일치 로 인한 정책 KL 발산 및 훈련 붕괴 문제 를 해결하는 것을 목표로 합니다.#Review#MoE#Reinforcement Learning#Training Stability#Routing#Policy Alignment#Rollout Routing Replay#LLMs2025년 10월 27일댓글 수 로딩 중