[논문리뷰] AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning본 논문은 장기 에이전트 학습(Long-horizon agentic RL)에서 기존의 sparse outcome reward 기반 학습이 가진 크레딧 할당의 불명확성을 해결하고자 합니다 .#Review#Agentic Reinforcement Learning#Self-Distillation#Credit Assignment#Recursive Belief Update#LLM#Policy Optimization2026년 8월 6일댓글 수 로딩 중