[논문리뷰] Multi-Turn On-Policy Distillation with Prefix Replay본 논문은 에이전트 기반 학습(Agentic Tasks)에서 Multi-Turn On-Policy Distillation (OPD)이 겪는 높은 비용과 효율성 문제를 해결하고자 합니다. 기존 OPD는 매 업데이트마다 학생 모델이 환경과 상호작용하고 교사 모델을 조회해야 하므로, 환경 배포 및 추론 비용이 매우 높습니다 .#Review#On-Policy Distillation#Knowledge Distillation#Multi-Turn#Agentic Tasks#Prefix Replay#Distribution Shift2026년 7월 23일댓글 수 로딩 중