[논문리뷰] Weak-to-Strong Generalization via Direct On-Policy Distillation본 논문은 대규모 언어 모델의 post-training 단계에서 발생하는 RLVR(Reinforcement Learning with Verifiable Rewards)의 높은 컴퓨팅 비용 문제를 해결하고자 합니다.#Review#Weak-to-Strong Generalization#Reinforcement Learning#On-Policy Distillation#Policy Shift#Implicit Reward#Post-Training#Large Language Models2026년 7월 13일댓글 수 로딩 중
[논문리뷰] Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals본 논문은 기존 LLM 사후 학습 방식이 탐색(exploration)과 분포 정렬(distribution alignment)을 강하게 결합하여 컴퓨팅 효율성과 확장성을 저해하는 문제를 해결합니다.#Review#Post-training#Proxy Exploration#Update Signal Transfer#LLM Alignment#Modular Training#Weak-to-Strong Generalization2026년 7월 13일댓글 수 로딩 중
[논문리뷰] DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes본 논문은 LLM의 추론 성능 향상을 위해 외부의 강력한 teacher 모델이나 복잡하게 큐레이션된 학습 데이터에 의존해야 하는 기존 RL 패러다임의 한계를 해결하고자 합니다. 기존 방식들은 학습 데이터의 품질이나 교사의 지식 수준에 따라 성능이 제약되는 structural limitation을 가지고 있습니다.#Review#Reinforcement Learning#Reasoning Models#Denoising Reasoning#Weak-to-Strong Generalization#Self-correction#Large Language Models2026년 5월 27일댓글 수 로딩 중