[논문리뷰] Calibrating Teacher--Student Discrepancy for On-Policy Distillation본 연구는 On-Policy Distillation에서 교사와 학생 간의 discrepancy가 교사 모델 자체의 변동성인 TSD를 포함하고 있어 효율적인 학습을 방해한다는 점을 문제로 제기합니다 .#Review#Knowledge Distillation#On-Policy Distillation#Teacher Self-Deviation#Calibration#Mathematical Reasoning#Privileged Information2026년 9월 20일댓글 수 로딩 중
[논문리뷰] What Does Privileged Information Add to On-Policy Self-Distillation?본 연구는 OPSD 과정에서 Teacher에게 제공되는 Privileged Information이 실제로 distillation 그 이상의 추가적인 학습 이득을 제공하는지, 그 기여도를 명확히 분리하고자 합니다.#Review#On-policy Self-distillation#Privileged Information#Language Model#Reasoning#AMPLE-Math#Knowledge Distillation#Training Dynamics2026년 9월 17일댓글 수 로딩 중
[논문리뷰] RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning본 논문은 Agentic task에서 사용되는 기존의 On-Policy Distillation (OPD) 방식이 가지는 고질적인 한계들을 해결하고자 합니다 .#Review#Agentic Reinforcement Learning#On-Policy Distillation#Privileged Information#Adaptive Retirement#LLM Agents#GRPO2026년 9월 17일댓글 수 로딩 중
[논문리뷰] One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation본 논문은 On-Policy Self-Distillation (OPSD) 기반의 학습 기법들이 직면한 collapse 문제를 심층적으로 분석하고 체계화합니다.#Review#On-Policy Self-Distillation#OPSD#Reinforcement Learning#Knowledge Distillation#Language Model#Collapse#Privileged Information2026년 9월 7일댓글 수 로딩 중
[논문리뷰] Best Practice Critic Optimization본 논문은 LLM RL 학습에서 Critic-based 접근 방식이 가지는 고질적인 불안정성과 성능 저하 문제를 해결하고자 한다. 기존의 Group-based methods는 여러 응답을 샘플링하여 비교함으로써 Critic 학습을 피하지만, 이는 높은 컴퓨팅 자원을 소모하며 토큰 레벨의 정밀한 학습을 제한한다.#Review#Reinforcement Learning#Large Language Models#Critic-Based Methods#Best Practice Critic Optimization#Generalized Advantage Estimation#Privileged Information2026년 8월 25일댓글 수 로딩 중
[논문리뷰] OPD-V: Visual On-Policy Self-Distillation with Modality Balance본 논문은 MLLM의 시각적 추론 성능을 향상시키는 On-Policy Self-Distillation (OPSD) 과정에서 발생하는 Modality Imbalance 문제를 해결하고자 합니다.#Review#Multimodal Large Language Models#On-Policy Self-Distillation#Modality Imbalance#Visual Reasoning#Privileged Information#Modality-Balance Trust Region2026년 8월 5일댓글 수 로딩 중
[논문리뷰] H^2SD: Hybrid Hindsight Self-Distillation본 논문은 기존의 RLVR 방식이 가지는 희소한 Supervision(Sparse supervision) 문제와 OPSD 및 RLSD 사이의 성능-안정성 트레이드오프 문제를 해결하고자 한다 .#Review#Reinforcement Learning#Self-Distillation#LLM Reasoning#Credit Assignment#Hybrid Hindsight#Privileged Information2026년 7월 21일댓글 수 로딩 중
[논문리뷰] dOPSD: On-Policy Self-Distillation for Diffusion Language Models본 논문은 dLLM의 추론 성능을 향상시키기 위한 효과적인 post-training 방법론의 부재 문제를 다룬다. 기존의 Supervised Fine-Tuning은 off-policy 문제로 인한 exposure bias에 취약하며, RLVR은 보상이 희소하고 sequence-level에 국한된다는 한계가 있다.#Review#Diffusion Language Models#On-Policy Self-Distillation#Privileged Information#Denoising Trajectory#Reasoning2026년 7월 6일댓글 수 로딩 중
[논문리뷰] DOPD: Dual On-policy Distillation본 논문은 OPD 환경에서 특권 정보를 주입할 때 발생하는 Privilege Illusion 문제를 해결하고자 합니다.#Review#On-policy Distillation#Privileged Information#Privilege Illusion#Advantage-aware#Dual Distillation#Large Language Model#Vision-Language Model2026년 6월 30일댓글 수 로딩 중
[논문리뷰] The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes본 연구는 OPD와 OPSD가 시스템 프롬프트 및 지식 내재화에는 효과적이나, 최근 연구들에서 보고된 학습 불안정성(instability) 및 성능 저하(degradation) 문제를 근본적으로 규명하고자 합니다.#Review#On-Policy Distillation#Self-Distillation#Language Models#Reverse-KL#Privileged Information#Optimization Stability#RLVR2026년 5월 12일댓글 수 로딩 중
[논문리뷰] Enhancing Object Detection with Privileged Information: A Model-Agnostic Teacher-Student Approach본 논문은 객체 탐지 성능을 향상시키기 위해 훈련 시에만 접근 가능한 특권 정보(Privileged Information, PI) 를 활용하는 LUPI(Learning Under Privileged Information) 패러다임을 통합하는 것을 목표로 합니다.#Review#Object Detection#Privileged Information#Teacher-Student Learning#Knowledge Distillation#Model-Agnostic#Bounding Box Masks#UAV-based Detection2026년 1월 8일댓글 수 로딩 중