[논문리뷰] Latent On-Policy Self-Distillation본 논문은 기존 OPSD 방식이 의존하는 수작업(hand-crafted) 방식의 privileged context가 모델의 확장성과 범용적 자가 진화(self-evolving)를 저해하는 핵심 병목임을 지적합니다.#Review#On-Policy Self-Distillation#Latent Context#Agent Evolution#Reinforcement Learning#End-to-End Learning#Privileged-Margin Constraint2026년 8월 16일댓글 수 로딩 중