[논문리뷰] Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It본 논문은 LLM의 RLVR(Reinforcement Learning with Verifiable Rewards) 환경에서 발생하는 심각한 Training-Inference Mismatch 문제를 해결하고자 합니다.#Review#Reinforcement Learning#LLM#Training-Inference Mismatch#Importance Sampling#Mixture-of-Experts#Calibrated Importance Sampling2026년 9월 28일댓글 수 로딩 중