[논문리뷰] TTPO: Test-Time Policy Optimization본 논문은 LLM의 test-time reasoning 능력을 향상시키기 위한 test-time training(TTT) 환경에서 ground-truth label 부재 문제를 해결합니다. 기존의 OPSD나 RLVR 기법은 정답을 필수로 요구하지만, test-time에는 이러한 정보를 얻을 수 없습니다.#Review#Test-Time Policy Optimization#Test-Time Training#On-Policy Self-Distillation#Reinforcement Learning#Chain-of-Thought#Mathematical Reasoning#Asymmetric Objective2026년 8월 27일댓글 수 로딩 중