[논문리뷰] On-Policy Delta Distillation for Multilingual Math Reasoning본 논문은 LLM의 수학적 추론 능력을 향상시키기 위한 On-Policy Distillation 기법을 다국어(한국어, 일본어 등) 환경으로 확장하는 것을 목표로 합니다. 기존 연구들은 주로 영어 중심의 추론 벤치마크에 집중되어 있어, 비영어권 언어에서의 OPD 및 OPD^2^ 성능 검증이 미흡한 실정입니다.#Review#On-Policy Distillation#OPD^2^#Multilingual Math Reasoning#LLM Post-training#Reasoning Transfer#English-Korean Performance Gap2026년 8월 6일댓글 수 로딩 중
[논문리뷰] Distilled Reinforcement Learning for LLM Post-training본 논문은 기존 RL과 OPD가 가진 한계를 해결하기 위해 Distilled RL을 제안합니다. RL은 결과 중심의 coarse-grained 보상을 사용하여 credit assignment 문제를 겪으며 새로운 지식 습득에 한계가 있습니다.#Review#LLM Post-training#Reinforcement Learning#Knowledge Distillation#On-Policy Distillation#Credit Assignment#Policy Gradient#Cross-family Distillation2026년 7월 20일댓글 수 로딩 중
[논문리뷰] On-Policy Delta Distillation본 논문은 기존의 On-Policy Distillation (OPD) 방식이 교사 모델의 전체 출력 분포를 모방하는 데 그쳐, 추론 능력 향상에 필수적인 핵심 학습 궤적을 충분히 전달하지 못한다는 문제를 제기합니다 .#Review#Knowledge Distillation#On-Policy Distillation#Reasoning Capability#Delta Signal#LLM Post-training#Reinforcement Learning2026년 7월 19일댓글 수 로딩 중
[논문리뷰] Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations본 논문은 최신 LLM post-training의 표준이 된 OPD의 학습 동역학이 여전히 불투명하다는 점을 지적한다. OPD는 때때로 성능 향상을 이끌지만, 많은 경우 불안정성을 보이거나 탐색 붕괴를 초래하며 심지어 outcome-based RL보다 성능이 저하되기도 한다 .#Review#On-Policy Distillation#LLM Post-training#Reinforcement Learning#Exploration Catalyst#Pathology#Signal Regulation2026년 7월 16일댓글 수 로딩 중
[논문리뷰] STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability본 논문은 RLVR 기반의 LLM 학습 과정에서 빈번하게 발생하는 Policy Entropy Collapse 문제를 해결하고자 합니다. 기존의 GRPO는 학습이 지속됨에 따라 출력 다양성이 사라지고 모델이 조기에 수렴하는 현상을 겪으며, 이는 장기적인 포스트 트레이닝의 병목 현상으로 작용합니다 .#Review#Reinforcement Learning#Policy Entropy#GRPO#Advantage Reweighting#Surprisal#LLM Post-training#Credit Assignment2026년 6월 17일댓글 수 로딩 중
[논문리뷰] Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation본 연구는 OPD가 일반적인 Supervised Fine-tuning(SFT)과 달리 어떤 기하학적 특성을 가지며, 왜 RLVR(Reinforcement Learning from Verifier-derived Rewards)과 유사한 sparse한 업데이트 양상을 보이는지 규명합니다.#Review#On-policy Distillation#Parameter Sparsity#Model Geometry#Subnetwork Masking#LLM Post-training#Optimizer Dynamics2026년 6월 14일댓글 수 로딩 중
[논문리뷰] Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders본 논문은 LLM post-training에서 데이터 엔지니어링이 모델 성능 향상의 핵심임에도 불구하고, 기존 방식들은 주로 외부 피드백(인간 선호도, 보상 모델, rollout 결과 등)에 의존하여 비용이 높고 효율성이 제한적이라는 문제에서 출발한다.#Review#Sparse Autoencoder#LLM Post-training#Reinforcement Learning#Data Engineering#Mechanistic Interpretability#Curriculum Learning#Data Selection2026년 5월 27일댓글 수 로딩 중
[논문리뷰] EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL본 논문은 Large Language Models (LLMs)에 tool-use capabilities를 부여하는 Agentic Reinforcement Learning (Agentic RL)이 겪는 두 가지 주요 bottleneck, 즉 scalable하고 robust한 executable environments의 부족과 implicit human reasoning을 포착하는 현실적인 training data의 희소성을…#Review#Agentic Reinforcement Learning#Tool-Use Agents#Environment Synthesis#Trajectory Generation#Dependency Graph#LLM Post-training2026년 5월 19일댓글 수 로딩 중
[논문리뷰] Near-Future Policy Optimization본 논문은 RLVR 과정에서 on-policy 탐색이 갖는 한계를 극복하고 최적의 보조 학습 신호를 확보하는 문제를 다룹니다.#Review#Reinforcement Learning#RLVR#Mixed-Policy#Trajectory Quality#Variance Cost#Self-Taught RL#LLM Post-training2026년 4월 22일댓글 수 로딩 중
[논문리뷰] Self-Distilled RLVR본 논문은 OPSD 가 훈련 초기에는 성능 향상을 보이나, 곧 정보 누출(Information Leakage)로 인해 성능이 저하되는 원인을 규명하고 이를 해결하고자 합니다.#Review#LLM Post-training#Reinforcement Learning#Self-Distillation#Information Asymmetry#Credit Assignment#RLVR2026년 4월 5일댓글 수 로딩 중
[논문리뷰] Revisiting On-Policy Distillation: Empirical Failure Modes and Simple FixesLarge Language Model (LLM)의 Post-training에 있어 On-policy Distillation (OPD)은 student-generated rollouts에 대한 teacher feedback을 활용하기 때문에 매력적이다.#Review#On-policy Distillation#LLM Post-training#Sampled-token OPD#Variance Reduction#Local Support Matching#Truncated Reverse-KL#Top-p Rollout Sampling#Special Token Masking2026년 3월 26일댓글 수 로딩 중
[논문리뷰] Hail to the Thief: Exploring Attacks and Defenses in Decentralised GRPO이 논문은 Large Language Models (LLMs) 의 후처리 훈련에 사용되는 분산형 Group Relative Policy Optimization (GRPO) 시스템의 보안 취약점을 탐구합니다.#Review#Decentralized RL#GRPO#LLM Post-training#Adversarial Attacks#Data Poisoning#Defense Mechanisms#In-context Attack#Out-of-context Attack2025년 11월 13일댓글 수 로딩 중